Pith. sign in

Paper Citation Record · LEDGER

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

As of 16 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.13768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13768 v3

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:29.861150Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:11:16.568415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:13:18.015262Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c895a0b-76a4-440d-8973-a3d0e758408f · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.242665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.242665Z digest=sha256:4fe630115cce3cb60e1d270d2059b695cfe38047320c8efae080074ffe991e56

Observation c98ec8b3-3062-4b4e-8bd9-9074a259a4ab · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Analysis of thompson sampling for the multi-armed bandit problem

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.249582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.249582Z digest=sha256:256c519f609baf04a41194eaef8725bd6ad3f332d65e8bd745b5c47c364424b9

Observation dbd2d4cd-c320-46f3-a782-489dfb43dccf · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Thompson sampling for contextual bandits with linear payoffs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.255372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.255372Z digest=sha256:48bc8ba7a3b9d9c0b5a312525de00f696914f2b4e59c1dd0286a4cc3c3175df4

Observation e1ed65f2-ecaa-4a58-9c66-3dbd4c185620 · outbound

This paper cites Optimal Best-Arm Identification in Bandits with Access to Offline Data.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Optimal Best-Arm Identification in Bandits with Access to Offline Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.266065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.266065Z digest=sha256:688e1b27f6ccd48d1dead9d4866880902dae396f84fb8cace17599247c7eb91b

Observation fb1ea420-bc2c-40a9-8b9f-0a2a64753cc7 · outbound

This paper cites Exploration--exploitation tradeoff using variance estimates in multi-armed bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Exploration--exploitation tradeoff using variance estimates in multi-armed bandits

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.274118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.274118Z digest=sha256:60cc17d64743caf8785c6aeb548b29cea601943dd4403e08c4745e703885390e

Observation 01e262d1-166e-40a0-9384-e84988a08afc · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:31.817957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.279590Z digest=sha256:bda2d3615aebb2d43036687fda46ea4561cce09e381870b8f0f2751eebf03f1d

Observation b9a12319-d135-458e-9167-e33af680928d · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax regret bounds for reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.287028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.287028Z digest=sha256:6ad831ac917392f4b68f98563b06e8a768d64cd4deaa318e667b3ba141016319

Observation f44891fb-e7aa-4523-b982-6c7d55e07e19 · outbound

This paper cites Stochastic linear bandits robust to adversarial attacks.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Stochastic linear bandits robust to adversarial attacks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.297657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.297657Z digest=sha256:f12c4d0e784a9d744fe469348881e914c9017bd64a84863c072f33664bda93c8

Observation f2bb5257-693c-4215-80da-4b2166df8393 · outbound

This paper cites Offline contextual bandits with overparameterized models.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline contextual bandits with overparameterized models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.757168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.307453Z digest=sha256:9b2bdc2dda7bf8b15ea0a52e865c305734ba851f15252870dee63b1667480315

Observation 6f739da6-b875-4918-a870-78e524e0581b · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.321056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.321056Z digest=sha256:883cf7e8a5ac210349f63c2fbdbf4ba41cd8e612db8242e687db4a0bed311c75

Observation 92622cb2-85f6-4433-b05d-c2d61d0a5f07 · outbound

This paper cites Kullback-leibler upper confidence bounds for optimal sequential allocation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Kullback-leibler upper confidence bounds for optimal sequential allocation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.715206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.327688Z digest=sha256:d66d618e0a0e079daeb5854588fda52215e6b16e56a49b268f91ce0d5a2f41a8

Observation a0f51a63-b927-4230-a41c-a09f73794017 · outbound

This paper cites The Elliptical Potential Lemma Revisited.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The Elliptical Potential Lemma Revisited

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.334478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.334478Z digest=sha256:45800e40e67fe104d0a644a0d3be5233097df0989bd1dbccb12622fbe4d20620

Observation e6ef10ea-4654-4bf6-880f-6913f90944ba · outbound

This paper cites An empirical evaluation of thompson sampling.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis An empirical evaluation of thompson sampling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.343000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.343000Z digest=sha256:2ebdc44ecf96871c6ff762edc22ad6cb2c1b2c3f61783984358e23841751f67c

Observation 49cc509e-74a7-455c-adfd-0ed56063cc20 · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Information-theoretic considerations in batch reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.350532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.350532Z digest=sha256:ef4e68c74230af791db7448c9162fbed61c82b9498d6a5bf31457bbaaeafb8b6

Observation c42f4de1-f0de-40c9-89d5-596efbef1efa · outbound

This paper cites Leveraging (biased) information: Multi-armed bandits with offline data.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Leveraging (biased) information: Multi-armed bandits with offline data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.359279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.359279Z digest=sha256:534e2f087f2736c3c7d520bf78b789c4c0fa0e29f1c027d72fa77e558f2dc05d

Observation 0d4484bb-ebcd-47bf-ac44-dc2c28e5298d · outbound

This paper cites Contextual bandits with linear payoff functions.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Contextual bandits with linear payoff functions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.367609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.367609Z digest=sha256:663f52af6600d985c5f4749e20bca7e3df082d113148ffd277abd8afbce43a25

Observation 4f273f8b-d01d-4bb3-8e7c-f73427d23de8 · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Stochastic linear optimization under bandit feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.623429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.375257Z digest=sha256:6cac380b218ef9e3aaff305aefb1138cdef493d66afffa565732c08394ba514a

Observation 04905289-db52-4afe-b818-25bbb45f9a4a · outbound

This paper cites Minimax-optimal off-policy evaluation with linear function approximation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax-optimal off-policy evaluation with linear function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.595939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.381472Z digest=sha256:ef81d6afd62876b7235d22a9f84e790e90cc252e35a4b89fbe4ee4d95be2ebac

Observation 010d2342-131d-49e0-8545-81ebdaab5c71 · outbound

This paper cites The kl-ucb algorithm for bounded stochastic bandits and beyond.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The kl-ucb algorithm for bounded stochastic bandits and beyond

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.570089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.388370Z digest=sha256:0ecb8d530aeafbaa17fa926e8011a7d15135954223504f2e4489e0ae47281b91

Observation 1eda491f-0c24-472c-ac81-c687a01adbb1 · outbound

This paper cites Guidelines for reinforcement learning in healthcare.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Guidelines for reinforcement learning in healthcare

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.534809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.396228Z digest=sha256:b7b055dade5a2160e998f5ab64ffeb5b35bcf13b38b6dac217b137b978e397c2

Observation 5ea1c3bb-61bf-45f5-b3c4-c629bc3c2ccc · outbound

This paper cites The movielens datasets: History and context.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The movielens datasets: History and context

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.404480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.404480Z digest=sha256:93f8c5e2da6a045719fa862e58bcc1eac5c38a50399a57decab5e791aa086632

Observation b265bfe8-4724-46cf-9f59-ffb10609395e · outbound

This paper cites A reduction from linear contextual bandits lower bounds to estimations lower bounds.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A reduction from linear contextual bandits lower bounds to estimations lower bounds

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.483752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.416232Z digest=sha256:06a7f536554e915ff5305fb21416625ec4906ae2f33cb068f35ffa72f36e10f7

Observation fe546f33-8415-4f92-993b-e61d5dfc2339 · outbound

This paper cites Deep q-learning from demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Deep q-learning from demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.425884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.425884Z digest=sha256:6ba8e73d4374d75d6c37bc0cb07ff9e5e43deeef3d238dbd6e875d9e37c9c331

Observation 4d80e72d-04bd-4be8-ab07-65ff2124aa24 · outbound

This paper cites Optimal best-arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Optimal best-arm identification in linear bandits

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.433662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.433662Z digest=sha256:b85548d9369db72ad2df740312cebcd8a6b540b41d21ff127d6d4cf9c2f482d2

Observation 582dc0a1-c169-4ace-9248-d63eb78cc5ae · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Provably efficient reinforcement learning with linear function approximation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.440235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.440235Z digest=sha256:4b7e0a7cb86effaf9392b3d59a32a1a72b8d76df8efc65793fe5bb24dae7a257

Observation 549a36c1-cb33-402e-a0db-ded1c5cc646a · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.398827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.447002Z digest=sha256:7a966e0e58c73e9fe3255820469a2b2dbc41ef2737bf5bcd4cb8e7e3436270f8

Observation eb141d2d-c21a-4f56-a2aa-89f723036659 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084--5096.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084--5096

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.453644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.453644Z digest=sha256:426bb8ab2ae7312634a445c2f8d147d13124bb42a263789bb5880a5d915aacdf

Observation e7ac0e9b-76de-4530-9ae9-4ef74f1c9216 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Conservative q-learning for offline reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.460288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.460288Z digest=sha256:fd5adf49a8db1fcff80d5abf756c6b74108d2d2bc46279d36225f7b96f094eaa

Observation 9dd65cf7-1b34-4045-80ca-b603d1c277c5 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Asymptotically efficient adaptive allocation rules

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.323438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.467738Z digest=sha256:72043d8d8a0f87815dfaf44921e44331ec6636fed2342ad1a0e3ab62dc461ef4

Observation 48770e6b-2a6e-4cbd-a145-19f174a6274f · outbound

This paper cites Bandit algorithms.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bandit algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.476237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.476237Z digest=sha256:7dda600520584291ae59e76c3518af5506d32e1bf5608db614490283da825eae

Observation 92c2d990-9539-41b7-8a02-2a9f06ad463e · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.483421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.483421Z digest=sha256:a74cc97f1ac8f101e84eb7d10a1aae471c25681459ab8bea3cc7a4d3067c2c33

Observation 1fd1f3f3-1d2e-4905-9663-0cde19d7c110 · outbound

This paper cites Lee, Yuejie Chi, and Yuxin Chen.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Lee, Yuejie Chi, and Yuxin Chen

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.258080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.490173Z digest=sha256:8a5d1897c214dc244d5c539d234405ab40abff341fd5599635e709e4bfc1b597

Observation 44abf8ef-5c8d-4d7b-bf1a-d4459351a26a · outbound

This paper cites Settling the sample complexity of model-based offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Settling the sample complexity of model-based offline reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.229101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.496105Z digest=sha256:b6a3fc51fa8b015ba48e4cce017df57bf96ec16b8ace6c207f3930174ce05250

Observation 6b53940f-e37c-4738-a6b0-c4f072530764 · outbound

This paper cites Pessimism for offline linear contextual bandits using l _p confidence sets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Pessimism for offline linear contextual bandits using l _p confidence sets

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.208203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.504425Z digest=sha256:80619a4a6731d3750d003337236fba6275f8f1e8d4a53820113caec431d99de4

Observation e61f07a1-673f-495b-88fb-876bcfaadb62 · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A contextual-bandit approach to personalized news article recommendation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.512902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.512902Z digest=sha256:aa4f5c7c3860ca76cb1b12b18e47404d855a3cefa6ce0465b10f4a6442879723

Observation 987369e1-50d9-462b-ab7c-2165fc6a5f6b · outbound

This paper cites Fast active learning for pure exploration in reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Fast active learning for pure exploration in reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.159599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.520932Z digest=sha256:115ed98549b957b0d4d04422a9a4bf90c3056d52cec0c5f8f9a7be2a1aad826a

Observation f557d306-35f1-454a-b0b5-d501c09e3a6d · outbound

This paper cites Efficient memory-based learning for robot control.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Efficient memory-based learning for robot control

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.530093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.530093Z digest=sha256:ef3b5dd697d781c3794089b638b12845d71cccd39a9b6530d23709fbae76ac4c

Observation 548bff10-8852-4329-83df-625b62a4ff98 · outbound

This paper cites Collaborative-filtering.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Collaborative-filtering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.105132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.539479Z digest=sha256:820ab31ccd29c8e735d1eecf0996c3dfbff34af8feb577ec047e33cd8165d3f2

Observation c9ac7849-878d-47fc-a732-d1156fe9ac2f · outbound

This paper cites Finite-time bounds for fitted value iteration.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Finite-time bounds for fitted value iteration

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.547938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.547938Z digest=sha256:d4d8dc2d9fc5c3868cbb4505052892df93b6bc35c2dd9a640a496922ef15a6f6

Observation 70166982-94ee-4140-930b-0fbe349720a1 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Overcoming exploration in reinforcement learning with demonstrations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.556309Z digest=sha256:3e83625db8534c23bbf8572f9df7ed5bd13714400d0a6d5a74dafc312980038b

Observation f668b8b1-7887-44cc-91b5-c1c598401423 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.565639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.565639Z digest=sha256:bd22b3c237dbd7b41e2a0a52431087b3c96bad03636679507dfe888a736faaf7

Observation a68c2aa5-0991-45f9-96ce-9013e9e23a41 · outbound

This paper cites Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.572118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.572118Z digest=sha256:6a319479d5f29b5e22755fd7ac25f8777eb8bd16b1bf424c3344895cfda3fc99

Observation 42d21d93-8346-44dc-a5be-0bd20480f589 · outbound

This paper cites Cutting to the chase with warm-start contextual bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Cutting to the chase with warm-start contextual bandits

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.008409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.579950Z digest=sha256:725915f53559f4407fb06814a6c3194daafeda251d7f62f4b69a9f559ead989e

Observation a23cd302-2197-4e6e-a6d0-ab54448f3761 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.587974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.587974Z digest=sha256:2a94165c77744534a4279040ebaeedced0f45578df242c13625169ab75fddf5b

Observation c2f7c428-45d9-4047-8262-7d8052ce2572 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.981176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.594940Z digest=sha256:be659a0bd3b649e69ea95636f75dd5206b9a91fb677ac2ab67d80324fc3c17e0

Observation 6c010a11-7c86-4337-8617-ef321ed93183 · outbound

This paper cites Agnostic System Identification for Model-Based Reinforcement Learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Agnostic System Identification for Model-Based Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.600345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.600345Z digest=sha256:aeb24d1f70d0f04f9797a741baefc4b0fdd0ab871b58d5a4f44ccf06d4487f48

Observation ef0b3f80-f1e1-4e99-989f-385b3fdab78b · outbound

This paper cites Bandits with Mean Bounds.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bandits with Mean Bounds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.612927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.612927Z digest=sha256:1bf7a94d741d83e4fadbbb69ed82189288bf1080f7793658a9425ac6b99dbc5d

Observation ec621217-6fcc-48d3-98c6-c099f076e4ee · outbound

This paper cites Multi-armed bandit problems with history.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Multi-armed bandit problems with history

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.946397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.622678Z digest=sha256:c83f6eb6acec258d004954a794fc8b5bbf0f2114f9b80cc4f8122e4395e3ff96

Observation 796f590c-eb96-45de-829c-327756193768 · outbound

This paper cites User cold-start problem in multi-armed bandits: When the first recommendations guide the user’s experience.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis User cold-start problem in multi-armed bandits: When the first recommendations guide the user’s experience

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.632051Z digest=sha256:ef3865acfbf175d9f7f6608c2c65a52a9f8990afd8b2f3eb856b0218fa303966

Observation 072ccc22-bb4b-4321-8cbd-af8035632d12 · outbound

This paper cites Best-arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Best-arm identification in linear bandits

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.869269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.648294Z digest=sha256:5e0e4cb1c59a1cb892765251fb3e2e9d8495e1db004216322e743bebb68f6a6b

Observation 905dd96b-4a67-4054-979a-a5edb08b2da2 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.656847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.656847Z digest=sha256:45e668cd70d66c8636c7a90edd146510c17fc1f28244038ecf328fe6ef68bb3d

Observation d557b812-f358-4906-9f25-cbfe7e9a1778 · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:30.840977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.669284Z digest=sha256:c2e40c20b7cf79a83f1e55719e711a8b5057b526ec8baf87fa4b7c980bbfeb92

Observation ef0ef3db-819e-4f9f-995d-2cdc6b32b605 · outbound

This paper cites Algorithms for reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Algorithms for reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.677081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.677081Z digest=sha256:7d95da13e720cc8e6b13f4d6f0d824d53177d7b0d2fc046d6aba6f7749aad9b1

Observation 4fabacf2-bb9e-4fca-b173-82e627b3c8fc · outbound

This paper cites A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:30.160097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.685128Z digest=sha256:b61f5603d62d28419109e701dd631239d429f21b239ff0b507b348a4c2d991d6

Observation 55fb2b55-bc45-4ca8-842e-fe744e94b712 · outbound

This paper cites Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.691317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.691317Z digest=sha256:6922fb609fe02cfc9ac0c9385c7375ebfcd30972afe952cbd9e6a03dad6b8760

Observation df602222-d459-41ae-9e10-d48e203e50b1 · outbound

This paper cites Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.799656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.697435Z digest=sha256:a30e3b89c6c4c0d8acd91bcb201c74de5447874ea5dec43b0f1ace0245a84ee0

Observation b0332b30-f19e-40ce-8a94-878cb945dace · outbound

This paper cites On the likelihood that one unknown probability exceeds another in view of the evidence of two samples.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.705219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.705219Z digest=sha256:8881d50d5e23b6a96a69a3524056226f15000ea03adcf4ddd31317a6d2f25724

Observation f3f4b2b8-416b-4af7-97f8-942f265df193 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.713695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.713695Z digest=sha256:2fd8231361d5dec2e462d03eb5a9f77cbe87220fb0d1e1fde84b09a47a9c04da

Observation a9c0dbe8-6e5e-4322-9442-971d1fa2aae3 · outbound

This paper cites Pessimistic model-based offline reinforcement learning under partial coverage.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Pessimistic model-based offline reinforcement learning under partial coverage

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.753059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.722442Z digest=sha256:6b17e4dfe4b49f154fc36aba190323e20592549eedde6a69c030d9b72a88bc19

Observation 3cac87ee-b04c-4e77-b05d-d5f45757dcd9 · outbound

This paper cites Representation Learning for Online and Offline RL in Low-rank MDPs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Representation Learning for Online and Offline RL in Low-rank MDPs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.732840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.732840Z digest=sha256:3759c3a7b47bec25a5e442d57a958b93083971b210d3273ebe5b32b3a6c9df92

Observation 87db7217-299f-4cca-b768-227f46ae0dfb · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.744641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.744641Z digest=sha256:cd80b2ff6e66bcc88c206bb337c58d3475c0bb166c0880d0cb4da8774c36ac80

Observation f508c3f2-438e-48e1-b5da-9d5d4cee3c12 · outbound

This paper cites Leveraging offline data in online reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Leveraging offline data in online reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.710969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.750323Z digest=sha256:9a4bb4777298103680c90e14bff5407d27a425788405f7c7cee451b67ae4102a

Observation 685e482b-1df0-4cd1-9d5c-2a4c6f90efbd · outbound

This paper cites Experimental design for regret minimization in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Experimental design for regret minimization in linear bandits

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.689024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.758386Z digest=sha256:f59c6fbd5f41f558cdb96e71e5772c194efc89e5594119d87f802b513348301c

Observation 09163372-f577-437b-8cb6-8605aac2909b · outbound

This paper cites Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:30.031698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.768398Z digest=sha256:b7c172a1c92df5b8edbf8949f9e0ca5dd09d5e0eac4682dc2dd299e1842d409b

Observation 0716a315-be0d-4d68-8c23-e9ece9cd99ac · outbound

This paper cites Bellman-consistent pessimism for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bellman-consistent pessimism for offline reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.665695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.778054Z digest=sha256:fb1bdc01cac158afc57ab3941677e09ba74dea9983aaec6c53dfcfba1e9ac8e2

Observation d779f11e-5fb4-4a27-b7e3-099c6753d832 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.642102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.787656Z digest=sha256:4a15431ea6b4b730d3a2edfabd98fc4610a4717601d96fd7db26c990d3e2f338

Observation ec161056-9f73-4450-8caf-ccb1b88a4a80 · outbound

This paper cites Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.794976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.794976Z digest=sha256:5b6a0383311c24af947a1b9e306c2033ad91d21987dd21a5eb1ae71ffb7d08ef

Observation e7181073-99bf-43f7-83f0-2b3174dbdbba · outbound

This paper cites Minimax optimal fixed-budget best arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax optimal fixed-budget best arm identification in linear bandits

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.801662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.801662Z digest=sha256:25b23e8fcc30b4e745a57decb42d08cb739469ea56770a6d9af087893358eb33

Observation 5d7e2902-3ce0-4eec-bcd5-fe73934058d7 · outbound

This paper cites Offline Reinforcement Learning for Wireless Network Optimization with Mixture Datasets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline Reinforcement Learning for Wireless Network Optimization with Mixture Datasets

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:29.968928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.810991Z digest=sha256:a0096da1c7b941222469ca45021dff2e160ae5dd6d93600a6eb3f6e8cdbb403b

Observation ff14bda5-a614-463d-96cc-aba9a321377b · outbound

This paper cites Provable benefits of actor-critic methods for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Provable benefits of actor-critic methods for offline reinforcement learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.605103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.819313Z digest=sha256:81ca18c549ad1c161657b8477d8891d520351b8d70bdcd1f7bc9b606284216ab

Observation 4af4ccfb-fb8b-4844-9c47-dd6400cd56b6 · outbound

This paper cites Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.829348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.829348Z digest=sha256:30b28c2fb05d4f0642688cd60cf3ef17d0e1eae66741c96bf004e87b425631bf

Observation c67f8696-c5ea-46ea-ab54-995e8465a604 · outbound

This paper cites @esa (Ref.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis @esa (Ref

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.840130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.840130Z digest=sha256:5955acd3a57f5c3b6a46de04d9c50d95181974672611aff6f9abb8f8aec69155

Observation cdcc5227-6b28-4a7b-9012-800c7f84a40a · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.850447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.850447Z digest=sha256:92287bdd3ed7cc40cbd6aba4010cc4d8dc591238d7d46fe8450ff38f71d17eb7

Observation 9079adcb-e6d0-4853-b3f4-ae45c8c1c524 · outbound

This paper cites page @startpage numbered @text Submitted to @long ( @short).

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis page @startpage numbered @text Submitted to @long ( @short)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.535409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.861150Z digest=sha256:c11c56adaf9f56ceec2c11f24c026723765ae08dbc367b971de457b85403d209

Pith citing papers

Observation c593b69d-45c1-400a-8621-a3891331fde4 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.017226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:9c8ec3217cbd4d379022d7f4b18cd15128efb638cc642d9cf50001232a7bbf31