Pith. sign in

Paper Citation Record · LEDGER

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

As of 24 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.13768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13768 v3

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:29.861150Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:11:16.568415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:13:18.015262Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c895a0b-76a4-440d-8973-a3d0e758408f · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.242665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.242665Z digest=sha256:e9102b5624eca766472209de2fa419c687b7f42bc111808f53755bfe2f3044bf

Observation c98ec8b3-3062-4b4e-8bd9-9074a259a4ab · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Analysis of thompson sampling for the multi-armed bandit problem

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.249582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.249582Z digest=sha256:01a4f5e6bc9445ad21ba700befcfd2bfa6464ea9ce68eac1c1b8dbc9bb1d814d

Observation dbd2d4cd-c320-46f3-a782-489dfb43dccf · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Thompson sampling for contextual bandits with linear payoffs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.255372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.255372Z digest=sha256:88ae02033853dbdce5f2dc2d94094bbcd7738b4a12d2bf692de68b446b9a9d82

Observation e1ed65f2-ecaa-4a58-9c66-3dbd4c185620 · outbound

This paper cites Optimal Best-Arm Identification in Bandits with Access to Offline Data.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Optimal Best-Arm Identification in Bandits with Access to Offline Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.266065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.266065Z digest=sha256:c632c9e14dd6ad3c6f8ac92617e67e81f478dc02d4e17adb2743fdb201e2697c

Observation fb1ea420-bc2c-40a9-8b9f-0a2a64753cc7 · outbound

This paper cites Exploration--exploitation tradeoff using variance estimates in multi-armed bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Exploration--exploitation tradeoff using variance estimates in multi-armed bandits

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.274118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.274118Z digest=sha256:625721385845ca8f0620bb4c7690f7b06092528b7cea3eb8f223f03d360a6818

Observation 01e262d1-166e-40a0-9384-e84988a08afc · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:31.817957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.279590Z digest=sha256:408e2f528cf2cca73d17060e1267d5eb17c98c729e66957e5cddcd137d0cf4eb

Observation b9a12319-d135-458e-9167-e33af680928d · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax regret bounds for reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.287028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.287028Z digest=sha256:bce5c575b774ec4926afe87c2b3b3bfdfcc6885878dbd24e48c7454bc8c6d153

Observation f44891fb-e7aa-4523-b982-6c7d55e07e19 · outbound

This paper cites Stochastic linear bandits robust to adversarial attacks.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Stochastic linear bandits robust to adversarial attacks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.297657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.297657Z digest=sha256:91da5513109cd1e3b53216310fd8ff1f7b6c1b22ed22ae725cf3b831ceaeaa9f

Observation f2bb5257-693c-4215-80da-4b2166df8393 · outbound

This paper cites Offline contextual bandits with overparameterized models.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline contextual bandits with overparameterized models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.757168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.307453Z digest=sha256:344e2b6e02a659d056bb65764c6122c781af412a01e2834b33352563b4e051e0

Observation 6f739da6-b875-4918-a870-78e524e0581b · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.321056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.321056Z digest=sha256:1567549d5842f902a7fe517b138245ff383fc2933ba142ddea9e1179eb01aaca

Observation 92622cb2-85f6-4433-b05d-c2d61d0a5f07 · outbound

This paper cites Kullback-leibler upper confidence bounds for optimal sequential allocation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Kullback-leibler upper confidence bounds for optimal sequential allocation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.715206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.327688Z digest=sha256:d6a3289e720591b93c20295142835a058ae1b3e2d6d87081d7c09c2b601e2982

Observation a0f51a63-b927-4230-a41c-a09f73794017 · outbound

This paper cites The Elliptical Potential Lemma Revisited.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The Elliptical Potential Lemma Revisited

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.334478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.334478Z digest=sha256:fbbaafa224e2e81894eb84803aeee21319c0ecd89cdf4ae28e211c2c7a3e8adc

Observation e6ef10ea-4654-4bf6-880f-6913f90944ba · outbound

This paper cites An empirical evaluation of thompson sampling.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis An empirical evaluation of thompson sampling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.343000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.343000Z digest=sha256:03370c3f121ebd1a9e893b8aa608ca7d00e177c03128707a5266f38e794fd255

Observation 49cc509e-74a7-455c-adfd-0ed56063cc20 · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Information-theoretic considerations in batch reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.350532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.350532Z digest=sha256:9fcf701938eb116553138c12b24a255505bff0ff6f6d0e3b17f89e957a18945f

Observation c42f4de1-f0de-40c9-89d5-596efbef1efa · outbound

This paper cites Leveraging (biased) information: Multi-armed bandits with offline data.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Leveraging (biased) information: Multi-armed bandits with offline data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.359279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.359279Z digest=sha256:6ef8522d93b1abbd925ba6ff1bef3455609a23b27517638ea76bf484c12e0a8b

Observation 0d4484bb-ebcd-47bf-ac44-dc2c28e5298d · outbound

This paper cites Contextual bandits with linear payoff functions.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Contextual bandits with linear payoff functions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.367609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.367609Z digest=sha256:8da7e5f0245e86bb0cb349e5500e3f2d2c465d81086ce84c7247d874eb472d29

Observation 4f273f8b-d01d-4bb3-8e7c-f73427d23de8 · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Stochastic linear optimization under bandit feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.623429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.375257Z digest=sha256:946d05e796532cde352f94ba95de6ce26dfbe4544ee3871a328bcd09a5f3b965

Observation 04905289-db52-4afe-b818-25bbb45f9a4a · outbound

This paper cites Minimax-optimal off-policy evaluation with linear function approximation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax-optimal off-policy evaluation with linear function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.595939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.381472Z digest=sha256:3606be840ac0f3bcf7fc4f99431e5db306176b7529e5d40713cc7aaff0d02d37

Observation 010d2342-131d-49e0-8545-81ebdaab5c71 · outbound

This paper cites The kl-ucb algorithm for bounded stochastic bandits and beyond.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The kl-ucb algorithm for bounded stochastic bandits and beyond

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.570089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.388370Z digest=sha256:864657ad0f7b07957f7e98f4379d7f916f2c6f2cbcd9421f38e34a2ed2c8d86e

Observation 1eda491f-0c24-472c-ac81-c687a01adbb1 · outbound

This paper cites Guidelines for reinforcement learning in healthcare.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Guidelines for reinforcement learning in healthcare

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.534809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.396228Z digest=sha256:307347a99a93d0342c7906af516fb887a40067fcb7f4feb2e715f35ed8b5c547

Observation 5ea1c3bb-61bf-45f5-b3c4-c629bc3c2ccc · outbound

This paper cites The movielens datasets: History and context.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The movielens datasets: History and context

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.404480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.404480Z digest=sha256:c5f0e0e23cd443e94b9c3b761da3b8e4e713e51e8e47d50bd16b2d144fd2ce2d

Observation b265bfe8-4724-46cf-9f59-ffb10609395e · outbound

This paper cites A reduction from linear contextual bandits lower bounds to estimations lower bounds.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A reduction from linear contextual bandits lower bounds to estimations lower bounds

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.483752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.416232Z digest=sha256:3478903a1281f447c6a84377f6ceacb58a6ad48af5cfc86474ed91df4cb6e294

Observation fe546f33-8415-4f92-993b-e61d5dfc2339 · outbound

This paper cites Deep q-learning from demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Deep q-learning from demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.425884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.425884Z digest=sha256:aca52fb6c73d8464c50ecd1bd1b83491cc7872c790e3435078ae460134e2aa4c

Observation 4d80e72d-04bd-4be8-ab07-65ff2124aa24 · outbound

This paper cites Optimal best-arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Optimal best-arm identification in linear bandits

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.433662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.433662Z digest=sha256:c1cacdaa3031dbdec5ede2b632ee0c4606ccf315ee2acee45da6f01d7030f5d1

Observation 582dc0a1-c169-4ace-9248-d63eb78cc5ae · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Provably efficient reinforcement learning with linear function approximation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.440235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.440235Z digest=sha256:e83b31316688dd6f7ca044bce823f5ff52594c60d5baeb5992c2b7cda0d72141

Observation 549a36c1-cb33-402e-a0db-ded1c5cc646a · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.398827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.447002Z digest=sha256:4a6af05924b293d178d394b010db651d85d3822d42c83f24f517a0c7087e46ac

Observation eb141d2d-c21a-4f56-a2aa-89f723036659 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084--5096.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084--5096

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.453644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.453644Z digest=sha256:4a0dc03726e22156c98082fbe652c1e9f28b29b63f0dc3bf97308511d6c58647

Observation e7ac0e9b-76de-4530-9ae9-4ef74f1c9216 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Conservative q-learning for offline reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.460288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.460288Z digest=sha256:3034e80f29015046fbe9609fb8ff725ba08d273f84a7b21e37fb13f615ea6308

Observation 9dd65cf7-1b34-4045-80ca-b603d1c277c5 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Asymptotically efficient adaptive allocation rules

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.323438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.467738Z digest=sha256:a805265757310fb3c200c2a151c68e8b93ea264dc19ba128a090cda97c6e9df2

Observation 48770e6b-2a6e-4cbd-a145-19f174a6274f · outbound

This paper cites Bandit algorithms.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bandit algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.476237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.476237Z digest=sha256:d373e9334bed7a183ed4f17574d325c086b499805d5558f07517132896c7bd4c

Observation 92c2d990-9539-41b7-8a02-2a9f06ad463e · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.483421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.483421Z digest=sha256:6b0af7e1e7756318232d62f7d06240e915e0603f27067b7ac85b298cba6b7643

Observation 1fd1f3f3-1d2e-4905-9663-0cde19d7c110 · outbound

This paper cites Lee, Yuejie Chi, and Yuxin Chen.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Lee, Yuejie Chi, and Yuxin Chen

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.258080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.490173Z digest=sha256:f1f0d2eba6c6a5091fbc001140e086d4b5745a0470358119e65d0886dc30ef5a

Observation 44abf8ef-5c8d-4d7b-bf1a-d4459351a26a · outbound

This paper cites Settling the sample complexity of model-based offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Settling the sample complexity of model-based offline reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.229101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.496105Z digest=sha256:6502fc8ee242f7f47dd74b2adb96a8439ab4e976ffbcb10bd541c6844b2ad9f8

Observation 6b53940f-e37c-4738-a6b0-c4f072530764 · outbound

This paper cites Pessimism for offline linear contextual bandits using l _p confidence sets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Pessimism for offline linear contextual bandits using l _p confidence sets

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.208203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.504425Z digest=sha256:d9e6153ad1e82d19bceeee9fbedfde8762d2dade5460e87697ae3986044726db

Observation e61f07a1-673f-495b-88fb-876bcfaadb62 · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A contextual-bandit approach to personalized news article recommendation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.512902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.512902Z digest=sha256:50fd6191a5dfedb036b764645e5c4a3b60e1f70e5c69313779a4f7bdae8928a6

Observation 987369e1-50d9-462b-ab7c-2165fc6a5f6b · outbound

This paper cites Fast active learning for pure exploration in reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Fast active learning for pure exploration in reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.159599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.520932Z digest=sha256:3d30af09810e54e04bc5b87b55275a39aa47bfc00f082ab59bad1ec80947e218

Observation f557d306-35f1-454a-b0b5-d501c09e3a6d · outbound

This paper cites Efficient memory-based learning for robot control.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Efficient memory-based learning for robot control

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.530093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.530093Z digest=sha256:475700160680cc4d4de28b8721975d2a3b431d86643afb3d17fc0a3bfe079616

Observation 548bff10-8852-4329-83df-625b62a4ff98 · outbound

This paper cites Collaborative-filtering.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Collaborative-filtering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.105132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.539479Z digest=sha256:d1d436e6ee640f08de9f741a7f8b529d385eb289884c0a9576b591e6b00d2aea

Observation c9ac7849-878d-47fc-a732-d1156fe9ac2f · outbound

This paper cites Finite-time bounds for fitted value iteration.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Finite-time bounds for fitted value iteration

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.547938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.547938Z digest=sha256:12f2bcdb4c9a82fe70cc5d4f7b385986af7c84de062be82ca111899082423878

Observation 70166982-94ee-4140-930b-0fbe349720a1 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Overcoming exploration in reinforcement learning with demonstrations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.556309Z digest=sha256:24e502be53f7e8141ea70e2ed11b4ff70de801d74d82e380b9138a50a59655bd

Observation f668b8b1-7887-44cc-91b5-c1c598401423 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.565639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.565639Z digest=sha256:62b00c2a1abe4d42de0ce9035d725b592f87aae3a6a69458944fa70fe3e4126e

Observation a68c2aa5-0991-45f9-96ce-9013e9e23a41 · outbound

This paper cites Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.572118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.572118Z digest=sha256:0bf589ba59a94d4ee101e283e58b40b0c9aac1a42e8fe1abd4a2fc4903d6a9d6

Observation 42d21d93-8346-44dc-a5be-0bd20480f589 · outbound

This paper cites Cutting to the chase with warm-start contextual bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Cutting to the chase with warm-start contextual bandits

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.008409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.579950Z digest=sha256:2517999401e127e8dceb5e613cfa8fdcca87db67fe89a2f9c00599afa08ea0bc

Observation a23cd302-2197-4e6e-a6d0-ab54448f3761 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.587974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.587974Z digest=sha256:509a3b8456ac88728efaf5148892a7338f778862bbd323cbcd09af5ee945b96c

Observation c2f7c428-45d9-4047-8262-7d8052ce2572 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.981176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.594940Z digest=sha256:5c992c9b0f5dddd7ec37238bd2da4d6d3abeae4d7656711ec644c6bb0e48dacd

Observation 6c010a11-7c86-4337-8617-ef321ed93183 · outbound

This paper cites Agnostic System Identification for Model-Based Reinforcement Learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Agnostic System Identification for Model-Based Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.600345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.600345Z digest=sha256:55fdffc1f3c4353d60792961f20b8d025e11867b4225ea29d84039d4dc212a85

Observation ef0b3f80-f1e1-4e99-989f-385b3fdab78b · outbound

This paper cites Bandits with Mean Bounds.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bandits with Mean Bounds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.612927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.612927Z digest=sha256:44c2a635408b985231dab4eae449bd99be55203f8e982755c086729d3cb42850

Observation ec621217-6fcc-48d3-98c6-c099f076e4ee · outbound

This paper cites Multi-armed bandit problems with history.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Multi-armed bandit problems with history

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.946397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.622678Z digest=sha256:426dfe4bde7085105d15d430dc4d5c6eb77c1e04a8fb549c83a0befa0df100b8

Observation 796f590c-eb96-45de-829c-327756193768 · outbound

This paper cites User cold-start problem in multi-armed bandits: When the first recommendations guide the user’s experience.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis User cold-start problem in multi-armed bandits: When the first recommendations guide the user’s experience

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.632051Z digest=sha256:3cd6530cd29fa566171f45438001817a55965124cc0ae330421e8442f22280fe

Observation 072ccc22-bb4b-4321-8cbd-af8035632d12 · outbound

This paper cites Best-arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Best-arm identification in linear bandits

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.869269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.648294Z digest=sha256:5bf6ad0bf82ac02b6e1527908def378f45234642f5b8856bbe73c53e9acaca5d

Observation 905dd96b-4a67-4054-979a-a5edb08b2da2 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.656847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.656847Z digest=sha256:1a3c0794180f198298c630fc15297f04f469a8705a6d30760fbe18eefa601cb9

Observation d557b812-f358-4906-9f25-cbfe7e9a1778 · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:30.840977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.669284Z digest=sha256:0a3611995714ddbcc5f77c42b93938403f19578b957c12d9fa24962bc8a763a9

Observation ef0ef3db-819e-4f9f-995d-2cdc6b32b605 · outbound

This paper cites Algorithms for reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Algorithms for reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.677081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.677081Z digest=sha256:b33e41c1c64f7b8e592a815dc305ae1fc3370eba37f33ad8298889eab1713caa

Observation 4fabacf2-bb9e-4fca-b173-82e627b3c8fc · outbound

This paper cites A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:30.160097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.685128Z digest=sha256:7202c24edebbd64790a4a3875205528c664eed3a9f74637fa8456974f85fee78

Observation 55fb2b55-bc45-4ca8-842e-fe744e94b712 · outbound

This paper cites Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.691317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.691317Z digest=sha256:6e3b3b33d394e9b952716e5ee7214c8dbc2e57022f6385c435b7b726737c1b01

Observation df602222-d459-41ae-9e10-d48e203e50b1 · outbound

This paper cites Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.799656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.697435Z digest=sha256:ff62be6db84073c459d88814f60302f1d9abab436506786a7bdc632d0f1bfa0e

Observation b0332b30-f19e-40ce-8a94-878cb945dace · outbound

This paper cites On the likelihood that one unknown probability exceeds another in view of the evidence of two samples.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.705219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.705219Z digest=sha256:e3792cd74b278cb47d8dbe42ae8a024108a5c65362a9bad0cb416b54745da738

Observation f3f4b2b8-416b-4af7-97f8-942f265df193 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.713695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.713695Z digest=sha256:56a2b691b0a40c2dbc99299640b005d3a9b690d7b0968e3beadd2baef2cbbef5

Observation a9c0dbe8-6e5e-4322-9442-971d1fa2aae3 · outbound

This paper cites Pessimistic model-based offline reinforcement learning under partial coverage.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Pessimistic model-based offline reinforcement learning under partial coverage

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.753059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.722442Z digest=sha256:bdf425c574492d394dd7126d6a0d4d6611f5f736737529b69ed0aa6e59661501

Observation 3cac87ee-b04c-4e77-b05d-d5f45757dcd9 · outbound

This paper cites Representation Learning for Online and Offline RL in Low-rank MDPs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Representation Learning for Online and Offline RL in Low-rank MDPs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.732840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.732840Z digest=sha256:69afcfcea297c69dcea57bf4d4442ca9df1f879ad3197fc65d9799ef65293d2f

Observation 87db7217-299f-4cca-b768-227f46ae0dfb · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.744641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.744641Z digest=sha256:cc74b3d47345db1bf21f6873857814e87339e95e659f1087f8ca6ce574657a58

Observation f508c3f2-438e-48e1-b5da-9d5d4cee3c12 · outbound

This paper cites Leveraging offline data in online reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Leveraging offline data in online reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.710969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.750323Z digest=sha256:ade37e97d75ee655af8d2cb53c450054ceb9b9b2b207194de9e7ed3e3266c181

Observation 685e482b-1df0-4cd1-9d5c-2a4c6f90efbd · outbound

This paper cites Experimental design for regret minimization in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Experimental design for regret minimization in linear bandits

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.689024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.758386Z digest=sha256:cc3c69eb6b7d82743626d6121c5494828a227f15d2ca029c0cbf8397bf01affe

Observation 09163372-f577-437b-8cb6-8605aac2909b · outbound

This paper cites Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:30.031698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.768398Z digest=sha256:4c2a96d10dd297bf81852b67ffe8ef29e822f341718f944323aecba080e830c1

Observation 0716a315-be0d-4d68-8c23-e9ece9cd99ac · outbound

This paper cites Bellman-consistent pessimism for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bellman-consistent pessimism for offline reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.665695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.778054Z digest=sha256:d4351366bd75b51da6e11522af419f8e95eb0a296295a547e68a0b816c990caa

Observation d779f11e-5fb4-4a27-b7e3-099c6753d832 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.642102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.787656Z digest=sha256:e3b7092bef295109d9c0d57e42cdf7453c2a7bef3ef7569c0e4d13d69a8ba252

Observation ec161056-9f73-4450-8caf-ccb1b88a4a80 · outbound

This paper cites Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.794976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.794976Z digest=sha256:f16012356227169e994b5a6c864e67839e460ec2507b08d37c7d0e5fe4d349b8

Observation e7181073-99bf-43f7-83f0-2b3174dbdbba · outbound

This paper cites Minimax optimal fixed-budget best arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax optimal fixed-budget best arm identification in linear bandits

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.801662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.801662Z digest=sha256:9d4810dffcfde897761d7940dd33c691bb87cd54769badf57cccbbc9077ca04b

Observation 5d7e2902-3ce0-4eec-bcd5-fe73934058d7 · outbound

This paper cites Offline Reinforcement Learning for Wireless Network Optimization with Mixture Datasets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline Reinforcement Learning for Wireless Network Optimization with Mixture Datasets

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:29.968928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.810991Z digest=sha256:72658c7a04e70b593915e4551474f63355260998498c176b8bafff9d800c42b2

Observation ff14bda5-a614-463d-96cc-aba9a321377b · outbound

This paper cites Provable benefits of actor-critic methods for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Provable benefits of actor-critic methods for offline reinforcement learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.605103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.819313Z digest=sha256:d925f340d64ca3ace3b951d7ca72d0eeaccd5504e1c8b475af471bfd635aa78c

Observation 4af4ccfb-fb8b-4844-9c47-dd6400cd56b6 · outbound

This paper cites Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.829348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.829348Z digest=sha256:b5426eefed6f47b03f017cc98342a27a228cb5ccb0b5e2fa12a3b3eb9d59d1a9

Observation c67f8696-c5ea-46ea-ab54-995e8465a604 · outbound

This paper cites @esa (Ref.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis @esa (Ref

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.840130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.840130Z digest=sha256:62cf4a6f1fc555487a2a96a7bd4bd64089c854ec2f9a3a813fc09cfe0ccfb877

Observation cdcc5227-6b28-4a7b-9012-800c7f84a40a · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.850447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.850447Z digest=sha256:6997a193bf5e2dbbe7e3bc511d6f12a00f80e2a185fc30d80816633b7e3d345a

Observation 9079adcb-e6d0-4853-b3f4-ae45c8c1c524 · outbound

This paper cites page @startpage numbered @text Submitted to @long ( @short).

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis page @startpage numbered @text Submitted to @long ( @short)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.535409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.861150Z digest=sha256:39809c06f8e6abb34e54cb4223f1f8e1081a079b94de8d5aa4008a60898d6cbe

Pith citing papers

Observation c593b69d-45c1-400a-8621-a3891331fde4 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.017226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:c97122d22199b41798c0960cd2cabeb2599e5739ee9c0224acbc3fe977dfc7e3