Pith. sign in

Paper Citation Record · LEDGER

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures

As of 23 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 0 inbound Pith citation observations for arXiv:2501.02089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02089 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:16.651877Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 109 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cd67070-287d-4075-be8a-2fa3de722aa9 · outbound

This paper cites Im- proved algorithms for linear stochastic bandits.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Im- proved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.208948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.208948Z digest=sha256:aa979b37ad78b9f03cd527ece3b331bacfac6f02774ff9a154f6a9c64e86ac26

Observation 1529a580-e3fc-49a1-886e-e55ef4ae0448 · outbound

This paper cites Model-based reinforcement learning with a generative model is minimax op- timal.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Model-based reinforcement learning with a generative model is minimax op- timal

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.214445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.214445Z digest=sha256:dc08ee1c6f074c4e7dab697c5267808e0709123c5156c1c239efa35f47e058ca

Observation d79c0843-8f33-40dc-b94f-c4a8d025be67 · outbound

This paper cites Degenerate nonlinear programming with a quadratic growth condition.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Degenerate nonlinear programming with a quadratic growth condition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.219509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.219509Z digest=sha256:aafdf4773bbd1b8946b255ade2b3c1cbf703bd8bbfd850984071d4934b6962c3

Observation c4dce0c3-ce49-4107-bfd8-61e49fb6134b · outbound

This paper cites Fitted q- iteration in continuous action-space mdps.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Fitted q- iteration in continuous action-space mdps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.225342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.225342Z digest=sha256:52b077fd20c03a25f1322665ad1d2b2105d26743d6d69fb80fca7d712d4849e2

Observation 1b85e5d2-4428-4f56-bda2-9ae3cc6a4408 · outbound

This paper cites Learning the target network in function space.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Learning the target network in function space

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.230182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.230182Z digest=sha256:c0ce5822802d735751344925ac6e4332a81b0ceb855d870e4cfdeed33bbc3fcd

Observation 95f88b01-9202-49e2-ad3c-3b1126b78f35 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem, 2002.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Finite-time analysis of the multiarmed bandit problem, 2002

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.234126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.234126Z digest=sha256:ed2bc5bcb3d1b27c5e0ac9e13e685408f76dcd6f6bf5abf1ff1560c6b94e0c30

Observation 4bfac92d-a9d1-4350-9b10-538a891de260 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax regret bounds for reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.238833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.238833Z digest=sha256:eb5aa6c23fbba5ec0757f8f2a625bb023edcc0d992dd4d0d2b73f2f5ec8a3608

Observation 612dbc69-42f1-4026-aac5-f9cb2c4542aa · outbound

This paper cites Prov- ably efficient q-learning with low switching cost.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Prov- ably efficient q-learning with low switching cost

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.242675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.242675Z digest=sha256:226531791a886c83303c4fb013877104fb887e6a18bea51416ab5746e6b91ebd

Observation 391d16ad-2c33-4164-bce9-9a283ade1ad1 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.247682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.247682Z digest=sha256:75db0caff22c24345ff2cd48dc1b86df5f0dc58446d042db6482f4c8b4902268

Observation 6fed46ad-9f26-4721-a0fd-d6a4296b182e · outbound

This paper cites Dynamic programming.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Dynamic programming

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.252535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.252535Z digest=sha256:cb9272421a03e1fd9a6505679b0cb5d7e45d5742e8772d3687bc137026c7a4f5

Observation 2d5bcf53-36f9-43ae-95c6-59bda6936c10 · outbound

This paper cites Online learning with switching costs and other adaptive adversaries.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Online learning with switching costs and other adaptive adversaries

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.257046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.257046Z digest=sha256:2264689030c34de065e75abaf3e7c3baa892f36cc75e23d529e9437464c2403a

Observation 579c1e47-5e27-49f6-b5c0-15309c726438 · outbound

This paper cites Information-theoretic considera- tions in batch reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Information-theoretic considera- tions in batch reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.261458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.261458Z digest=sha256:22198ff2569fa23d07c7d0a9020a59f788e4cd62a4d3013078fa6b59ed42db11

Observation 6e30b8fe-222f-4a3d-8460-5f0773f20b1e · outbound

This paper cites Deep reinforcement learning from human preferences.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Deep reinforcement learning from human preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.265753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.265753Z digest=sha256:f86d50d0ab96b19c710970be6adc5f8204c11094f1cbfe9b658cccb973bd6e8d

Observation b77f5264-9dea-4525-966d-53eb0a2c6bb8 · outbound

This paper cites Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.269952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.269952Z digest=sha256:12128eec3fbb6107796e6d2c75c39c2c69aef27ce97984160b5602247b79c1df

Observation 30728537-2956-4a7d-9cae-85becbeb2052 · outbound

This paper cites Minimax-optimal off- policy evaluation with linear function approximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax-optimal off- policy evaluation with linear function approximation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.274588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.274588Z digest=sha256:aa767fca57be1628f43799852d4023db16026bc76ececebefb5b991b7c5ca4f5

Observation 2f664f1a-61d3-4d8b-b260-dafc2f2b3c11 · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Doubly Robust Policy Evaluation and Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.279147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.279147Z digest=sha256:6cfb5e1d93c0473370e1dbd28134ea669c4067e1307e53b8e4b24a76a407a83a

Observation 9ff4366c-7ce2-4045-b441-874bd0c7ad4b · outbound

This paper cites Tree-based batch mode reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Tree-based batch mode reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.284580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.284580Z digest=sha256:f76401b2635f585557d3fd741de75951347f61dd4bfd6d1fa1b896c11be38a9a

Observation 3e421dd0-e1f5-4858-aff6-0396b456a2f3 · outbound

This paper cites Discovering faster matrix multipli- cation algorithms with reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Discovering faster matrix multipli- cation algorithms with reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.289795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.289795Z digest=sha256:f2603e6e95baf55dfed03ad889417b7c332729d1bbf2f00f69a46eda3f5edcd3

Observation 1547a81a-a1db-4e9e-8690-9deba09ac06b · outbound

This paper cites Theory of statistical estimation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Theory of statistical estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.294226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.294226Z digest=sha256:6f48e00a4e2aa636c292050e47e5edbd6fba37fcfe1a793280e3b3435237e949

Observation b23adf15-01b8-408c-ba93-c0a3705e9685 · outbound

This paper cites A Provably Efficient Algorithm for Linear Markov Decision Process with Low Switching Cost.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures A Provably Efficient Algorithm for Linear Markov Decision Process with Low Switching Cost

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.298033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.298033Z digest=sha256:ee18aed20e32ddc2bb8985d2486dff9ff54270a5589d404d293e49788674cda0

Observation 6db73a91-af48-4e56-80e4-676867e272e1 · outbound

This paper cites Batched multi-armed bandits problem.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Batched multi-armed bandits problem

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.302003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.302003Z digest=sha256:13fc9625a7baa6a42aa372ed83766128bb72d3d8455cc070dfc819ccc043552f

Observation cac3d414-6306-424a-b5bb-2427cbdfb052 · outbound

This paper cites Off-policy deep rein- forcement learning by bootstrapping the covariate shift.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Off-policy deep rein- forcement learning by bootstrapping the covariate shift

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.305987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.305987Z digest=sha256:b0d350866eb686bc56454fd5a39c9b792be1194c5dd92ce39a40c75967da6dbf

Observation 961f3efe-270f-40b7-9217-70dd5c1cf63b · outbound

This paper cites Minimax pac bounds on the sample complexity of rein- forcement learning with a generative model.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax pac bounds on the sample complexity of rein- forcement learning with a generative model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.311486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.311486Z digest=sha256:b864872b91705968de38555c6a33f935c342d90ba9178297b75e656ae0c11bbf

Observation 689b4757-97cd-4604-b07c-f3f151f94402 · outbound

This paper cites Maxmin expected utility with non-unique prior.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Maxmin expected utility with non-unique prior

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.317103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.317103Z digest=sha256:8824867df061a0db2411b22f442a487a9ea7be00b3a56b75d1ee814bf7707c66

Observation 1b6a0a7d-db16-4d7a-98c6-b9ab21630074 · outbound

This paper cites Approximate solutions to Markov decision processes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Approximate solutions to Markov decision processes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.321994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.321994Z digest=sha256:13d12a5bbd0134476e8a2babefbcf749dc73e6dc485487103edb07ca3e981a4d

Observation 2fe80bfc-4112-4f5f-b37b-55194084191d · outbound

This paper cites Networkgym: Reinforcement learn- ing environments for multi-access traffic management in net- work simulation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Networkgym: Reinforcement learn- ing environments for multi-access traffic management in net- work simulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.326694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.326694Z digest=sha256:324296392af8047a69450477a7c9a9ee819dfb75cccb7f8bea5fc02da1649369

Observation 60ca303f-2507-4a8c-a827-34d6d588f2ca · outbound

This paper cites Consistent on-line off-policy evaluation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Consistent on-line off-policy evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.331632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.331632Z digest=sha256:48e875d1353134978842e657d685e8ddf3939b744f43e3f7e05b7a7cc5e1ac19

Observation 05eb0f54-0dd3-451b-b852-3dffe9027cc4 · outbound

This paper cites Bootstrapping fitted q-evaluation for off- policy inference.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bootstrapping fitted q-evaluation for off- policy inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.336214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.336214Z digest=sha256:56cab344dd97ef3e5275a9cff9872971d61914304f75bb81cd2dde594ce843ff

Observation 667d8a66-1ef5-4361-8463-37b978db587a · outbound

This paper cites Effi- cient estimation of average treatment effects using the estimated propensity score.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Effi- cient estimation of average treatment effects using the estimated propensity score

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.340340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.340340Z digest=sha256:be4217e68766c5fb5df7eb0019d8550312a9844d53f39cd838226e276f90bbd3

Observation d26e45f8-545a-4681-901f-2b34b8f84dab · outbound

This paper cites A generalization of sampling without replacement from a finite universe.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures A generalization of sampling without replacement from a finite universe

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.344523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.344523Z digest=sha256:c82baa2e7ea7c0cf6eebaad4de1ab0cf047afbb283315c9333e8e28b1a426433

Observation f0b9d2f9-6be9-4d6c-b1c6-7370f8df9449 · outbound

This paper cites Towards deployment-efficient reinforcement learning: Lower bound and optimality.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Towards deployment-efficient reinforcement learning: Lower bound and optimality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.348936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.348936Z digest=sha256:def795975d65edbd388d9ba7b0fd8a75b197f73d9891cbd7c4e1e3a77c759867

Observation 043356c3-92bf-4056-b959-a07c8de46a47 · outbound

This paper cites Aleatoric and epis- temic uncertainty in machine learning: An introduction to con- cepts and methods.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Aleatoric and epis- temic uncertainty in machine learning: An introduction to con- cepts and methods

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.353131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.353131Z digest=sha256:ddb36ddb075ef505273c7248479d63f21c29bcd46f4c899b6b6fc449d8f8976e

Observation 4c4b3a65-6f0c-404d-aa1f-1ff225bc51bd · outbound

This paper cites Doubly robust off-policy value eval- uation for reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Doubly robust off-policy value eval- uation for reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.788348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.357051Z digest=sha256:23faa28d077064460370ac02c6da4baf811aed3b9383f32bd776d38fe6a60e6e

Observation e6085e59-902e-41c4-91d9-e09f3fc8ff07 · outbound

This paper cites Offline reinforcement learning in large state spaces: Algorithms and guarantees.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline reinforcement learning in large state spaces: Algorithms and guarantees

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.775309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.360774Z digest=sha256:1ddcf941c809f4c69c20e1912d5a990b9180b17ab2d8da0758625bf6c74943a1

Observation f52c8254-9bed-41b8-9dff-eb535d10df04 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.761998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.365059Z digest=sha256:65eef4760e8846380ef2144209aa6478cf4b701a73ba1a3bc61bea9c39fd61db

Observation b8a3ab84-a748-4bcd-b827-85ff45cac308 · outbound

This paper cites Double reinforcement learning for efficient off-policy evaluation in markov decision processes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Double reinforcement learning for efficient off-policy evaluation in markov decision processes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.748834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.369711Z digest=sha256:96ab98417d69cfcda0d110b20d9da1f4f239c6edf41c54d0d44b1ba6d4bc42f5

Observation 5a83a642-146b-4251-9a99-64b8b41d7be2 · outbound

This paper cites Efficiently breaking the curse of horizon in off-policy evaluation with double reinforce- ment learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Efficiently breaking the curse of horizon in off-policy evaluation with double reinforce- ment learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.735715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.373992Z digest=sha256:2b97fe2cb8ed72dfcd282da272e11e457e2355bd5933b7a30399a34cd362aafe

Observation bafe5139-e632-4f6d-8cb9-5edc5238010c · outbound

This paper cites Near-optimal reinforce- ment learning in polynomial time.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal reinforce- ment learning in polynomial time

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.722332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.378249Z digest=sha256:129dc88db30967b3cc1f66a75be73494b503e02593e86c326b37a6de3e332baa

Observation 6226483e-26dc-4d44-9f62-557c512b6647 · outbound

This paper cites Introduction to empirical processes and semiparametric inference, volume 61.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Introduction to empirical processes and semiparametric inference, volume 61

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.710137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.383340Z digest=sha256:d12dc844b128bfc44ab00b76cafbf9965ec586553ae92c36200c73855a1615ac

Observation 8641d3d2-57e5-4f99-8e8a-5d4477510b90 · outbound

This paper cites Offline reinforcement learning with fisher divergence critic regularization.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline reinforcement learning with fisher divergence critic regularization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.698074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.387687Z digest=sha256:d7049d347dd3cff936f0351a92947da5076460b7953b62bcd2d09a30c3a70351

Observation 0663e18c-acf2-4cba-9ed9-b16a229433a0 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Conservative q-learning for offline reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.685186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.391784Z digest=sha256:e597179072eb178c4cfa92d508734aab19c2790a0d66e9e1245717923bc6622b

Observation 3aa17862-1c9b-4384-92c8-b34eecdbfe0a · outbound

This paper cites Minimax theory.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax theory

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.671509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.395913Z digest=sha256:ae412f80395153db2e20526df62731afa62ad40c09f031e8d09624ca375419a1

Observation 9a2ba5fa-fec8-4916-8dc0-a6bd09c3fd44 · outbound

This paper cites Batch policy learning under constraints.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Batch policy learning under constraints

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.658549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.400745Z digest=sha256:92a19888fe657d6f741ab3739823736219bdd71e98bcfbc9aac46542348e9167

Observation 53fbf41f-b39e-4273-baa9-63ed9b715275 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.405537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.405537Z digest=sha256:1cc262a6ade26ac29c521d05bae2cb3e70096edbf419133ceaad987a34cb257f

Observation 215ce0da-c7fa-4b64-8547-77caf4e57389 · outbound

This paper cites Breaking the sample size barrier in model-based reinforcement learning with a generative model.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Breaking the sample size barrier in model-based reinforcement learning with a generative model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.645299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.410690Z digest=sha256:688f8b366cfd26afef9337bbba47bc39a0a36f0fbe4269579c2964942f497809

Observation 49255547-06f2-4497-9765-454caa65547e · outbound

This paper cites Offline reinforcement learning with closed-form policy improvement operators.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline reinforcement learning with closed-form policy improvement operators

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.632786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.415116Z digest=sha256:1a3334a9b59ee1541429a1164147ede8ae5602f1f20edc9d2ec8fddb181c5a69

Observation f09321f7-cc91-43c8-804d-e6d4873c4936 · outbound

This paper cites Unbi- ased offline evaluation of contextual-bandit-based news article recommendation algorithms.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Unbi- ased offline evaluation of contextual-bandit-based news article recommendation algorithms

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.620627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.419615Z digest=sha256:e16e28531aa2347e59b7ba3e9c2f2a61e7eb42f5cf8512763bbfacdff59eb170

Observation 9a16b704-7ff1-4654-80e9-82d0632306c6 · outbound

This paper cites Monte Carlo strategies in scientific computing, volume 10.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Monte Carlo strategies in scientific computing, volume 10

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.608305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.423566Z digest=sha256:9d0b3425f9e31bea000efc871451aa10629446d9314fb300ff45776973a08a91

Observation 78954958-2c3a-4b48-8a42-23141e5ff526 · outbound

This paper cites Breaking the curse of horizon: Infinite-horizon off-policy es- timation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Breaking the curse of horizon: Infinite-horizon off-policy es- timation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.595029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.427799Z digest=sha256:9a00c45ff5b9553e5762840bd69f15db16656e11dbee58ad6aca603a5e9757b5

Observation 57a75e49-66aa-434d-9959-297a5db9a7af · outbound

This paper cites Off-policy policy gradient with stationary distribution cor- rection.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Off-policy policy gradient with stationary distribution cor- rection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.581997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.432172Z digest=sha256:17b6fb55d9e80fb3ae35836484d68f406e98ee00b90e53bf6cd1916734668e24

Observation ae97e546-50b3-4c9b-bf78-01671a0aa7ee · outbound

This paper cites Provably good batch off-policy reinforcement learning without great exploration.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Provably good batch off-policy reinforcement learning without great exploration

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.567624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.436590Z digest=sha256:6d92460a04c6e04649c99f5e529deb3e2053a0b776c1d3da3977e22b8ee4efc2

Observation 405675f4-ba25-4549-8375-532ad9945af0 · outbound

This paper cites Mildly conservative q-learning for offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Mildly conservative q-learning for offline reinforcement learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.554612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.441947Z digest=sha256:c9e782950f10d231aedab33b33bcebe38b2604357532c31a01225276b172fdde

Observation 9e7b0ac0-cd9f-41d8-acd5-3af09958e87e · outbound

This paper cites the distribution-norm to the res- cue.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures the distribution-norm to the res- cue

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.541828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.446115Z digest=sha256:a971747b4f4995950ffa69ca829bbf1c51eb60795be25c86e59af3e9d4310f17

Observation 54ed687c-d2c2-46b3-a236-996f74795263 · outbound

This paper cites Faster sorting algorithms discovered using deep reinforcement learn- ing.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Faster sorting algorithms discovered using deep reinforcement learn- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.529395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.450397Z digest=sha256:63cc50c9fc0ab520dcc8f45efd7a2317db772f7f65b3a49378ebb9997d4e99e6

Observation 1e0d4ece-f616-4e4f-81d8-bdb7858b2412 · outbound

This paper cites Neural adaptive video streaming with pensieve.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Neural adaptive video streaming with pensieve

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.516072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.455122Z digest=sha256:a4feddb1f59396a5bbb0a65b4daea1e4311a69f8e1eb04887cc9604a71846124

Observation 8b2e305d-de76-487c-835f-3755f4a0be3a · outbound

This paper cites Deployment-efficient reinforce- ment learning via model-based offline optimization.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Deployment-efficient reinforce- ment learning via model-based offline optimization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.502046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.459496Z digest=sha256:32744c6952cec85ee420e6fb18c006f4522370332256c487417bad8ea0505939

Observation cb40f6e4-1f27-4e42-bff5-f758f1b319b9 · outbound

This paper cites Dependent central limit theorems and in- variance principles.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Dependent central limit theorems and in- variance principles

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.487751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.464018Z digest=sha256:e4097046038c3c20c6a19ded4fa6534fdcfc46ea47918cd8f8dd4041d54d8eff

Observation 58ccbf41-73d4-412d-89f5-14b877a01c8b · outbound

This paper cites Variance-aware off-policy evaluation with linear function ap- proximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Variance-aware off-policy evaluation with linear function ap- proximation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.475680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.468042Z digest=sha256:b972f16f2d06c938bc01917a12ca59d056cc3e02887a554a7684465b5061111a

Observation e34d18aa-ef27-4d2a-9fb0-74054f8e5149 · outbound

This paper cites Human-level control through deep reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Human-level control through deep reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.462963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.471469Z digest=sha256:2d134b0f3585cbb0d92ac9e3b9c81d44804a1ec85f3301487d61e69b9ad019b9

Observation dd6bb0c5-0cfa-427b-b9ca-0fed1a73afff · outbound

This paper cites Bootstrapping: A nonparametric approach to statistical infer- ence.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bootstrapping: A nonparametric approach to statistical infer- ence

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.450599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.475970Z digest=sha256:378da0d3901798af352271da20ad40fb8c12952cd84aee88f129b4fa0bff8890

Observation e59f5fa6-37ca-4640-8f65-b706af237290 · outbound

This paper cites Finite-time bounds for fit- ted value iteration.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Finite-time bounds for fit- ted value iteration

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.437575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.479912Z digest=sha256:82073aacd770e527cd7bdf7cfd1dcff2d66ea5abd2df1cb5a8cd965a9f362e0e

Observation 193c6ec1-b31e-4457-95f6-cac07abce26a · outbound

This paper cites Marginal mean models for dynamic regimes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Marginal mean models for dynamic regimes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.422478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.483698Z digest=sha256:7145747eeda3da7f916be30bdbb6a7017d328c9a53bb51af206031c0f49aaa80

Observation 3ccf2f67-d769-44e9-9bce-ee8a2068faaf · outbound

This paper cites Dualdice: Behavior-agnostic estimation of discounted stationary distribu- tion corrections.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Dualdice: Behavior-agnostic estimation of discounted stationary distribu- tion corrections

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.409007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.487956Z digest=sha256:fe376d590eac8a3d4fc1ae58416af34abfc0d4ec092f4de754675d6f09145daf

Observation 95de1a03-1199-4c28-ad5b-2985258fee37 · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.491915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.491915Z digest=sha256:61b465d5c621265554294e550d24878fc08810d5968c9e640db6b8af3dc11492

Observation b0d4c04b-e228-4175-ae87-a8383b436721 · outbound

This paper cites Optimal medication dosing from suboptimal clinical ex- amples: A deep reinforcement learning approach.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Optimal medication dosing from suboptimal clinical ex- amples: A deep reinforcement learning approach

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.394785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.497573Z digest=sha256:a784cc048b7c1627fc9af4126a3342c9b3db085d5cc22be9ca6ab71858624f07

Observation c3585e0a-0abb-469f-9e0b-71f14dfd013e · outbound

This paper cites On sample-efficient of- fline reinforcement learning: Data diversity, posterior sampling and beyond.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On sample-efficient of- fline reinforcement learning: Data diversity, posterior sampling and beyond

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.381978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.502392Z digest=sha256:809c579f7ffa7ae6e4d90046f7021ba772fdb1f170708545b85b1142606f54df

Observation 19b323a4-cbe9-4848-895f-91a27aa5631a · outbound

This paper cites On instance-dependent bounds for offline reinforcement learning with linear function approximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On instance-dependent bounds for offline reinforcement learning with linear function approximation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.368016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.507044Z digest=sha256:8a6516a110a7881bf41414e74aa0c61e1b27b62a7d5da06d47af72d92db1841f

Observation 2d4b2c17-a3f3-47bd-bd3d-27744fb6f1b5 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.354576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.511379Z digest=sha256:1823a6b9bf5c51854aa03a91a0b3c97c898a487e8e42710298c1f0dc90ea87f0

Observation 2e04c9c0-c068-49f1-81b0-e36439b6e035 · outbound

This paper cites Batched bandit problems.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Batched bandit problems

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.341189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.516827Z digest=sha256:d810322b742d84432da1b098b550b6fcc26b452151a1aee4cb6240f774e016a8

Observation 4c44f74f-5eb9-4085-b8d9-9d5dd9f2e6d4 · outbound

This paper cites Approximate Dynamic Programming: Solv- ing the curses of dimensionality , volume 703.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Approximate Dynamic Programming: Solv- ing the curses of dimensionality , volume 703

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.326945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.521801Z digest=sha256:01999111bba75836c4ba4fde785eef88ba842e2f45bd43a53d99487c579e2058

Observation e5b586e5-8f70-4316-b061-bf6d2bf8990a · outbound

This paper cites Eligibility traces for off-policy policy evalua- tion.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Eligibility traces for off-policy policy evalua- tion

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.311762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.525885Z digest=sha256:42d1aad2ff9d79fb4ecd53d99ac6548cae494733db07940250ca31ab91e22e4f

Observation 45d70d9d-3624-4ccc-8f27-5dd679f15df5 · outbound

This paper cites Markov decision processes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Markov decision processes

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.298253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.529327Z digest=sha256:92352eaa0af71de1514acf191a57cd003c05965ae9e277dd42e67b43e2924149

Observation 563ed0ae-d2b1-41ed-91ff-088985f8c4ae · outbound

This paper cites Near-optimal deployment effi- ciency in reward-free reinforcement learning with linear func- tion approximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal deployment effi- ciency in reward-free reinforcement learning with linear func- tion approximation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.285372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.533749Z digest=sha256:f6c878a1d11576d2fbf99a96c4e3e1aecd7867ad6db23e5bd84c4f9aa27b5623

Observation 5cc730c1-b9ae-4717-bff6-d5aed1d48b13 · outbound

This paper cites Sample- efficient reinforcement learning with loglog (t) switching cost.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Sample- efficient reinforcement learning with loglog (t) switching cost

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.270273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.539089Z digest=sha256:c57d9984cfe323fb7e60e211cada66c552c240bc050adc1624ea97e96ccbf7b3

Observation 90471ee9-5843-4fbd-be07-5c0ef964fc56 · outbound

This paper cites Logarithmic switch- ing cost in reinforcement learning beyond linear mdps.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Logarithmic switch- ing cost in reinforcement learning beyond linear mdps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.255721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.543674Z digest=sha256:935d8ce839baff83f424b7b0678331b674716a5242d93cbdafda7c2d44bcc289

Observation d74fba3d-d6a5-40ca-bb2e-ac5dc47a8672 · outbound

This paper cites Bridging offline reinforcement learning and im- itation learning: A tale of pessimism.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bridging offline reinforcement learning and im- itation learning: A tale of pessimism

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.548582Z digest=sha256:58597a682c40d9a289cd324b0e6f0da7ddc630dac189d8090f318664ea281729

Observation 08b9b6dd-f4fb-4a25-bcee-9eba6b35c3e6 · outbound

This paper cites Nearly horizon-free offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Nearly horizon-free offline reinforcement learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.228950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.553237Z digest=sha256:5335f59363c731d8845e6b5fd1b2ef6e2c7a9bec0f8d16b9e152e00cc6ee5f71

Observation 73f5d59e-d8e3-46ff-8735-08a530c69867 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Mastering the game of go with deep neural networks and tree search

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.215211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.558147Z digest=sha256:54db6ac4b11e227eb9a67a8cda560610e3fcbd95820a78dcf056a561e3b76e3f

Observation 60a06f68-e025-467a-9587-758852db3281 · outbound

This paper cites Mastering the game of go without human knowledge.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Mastering the game of go without human knowledge

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.563181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.563181Z digest=sha256:7e13b736b7b1bb366c1eac1dbeef21dd8b0aa4ba05837b9e55e6f7e8be3dca96

Observation aa648aef-6fb2-4b3c-9e69-3af10e8a56ca · outbound

This paper cites Learning to summarize with human feed- back.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Learning to summarize with human feed- back

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.192685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.567487Z digest=sha256:3aa4acf64ae5625b30d5e226f8aa678de387e5ca5f20715a69edf3e224d116a4

Observation d0e89996-56c5-4de9-b783-0482330b2bde · outbound

This paper cites Reinforcement learn- ing: An introduction.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Reinforcement learn- ing: An introduction

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.179346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.571125Z digest=sha256:33a3c68ce8df674e72b47fd192db7243c3678a62ebc2d81271f95cac6a2afb2b

Observation 5c4d9419-5927-4a22-adc4-ac8db1fc9fd2 · outbound

This paper cites Finite time bounds for sampling based fitted value iteration.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Finite time bounds for sampling based fitted value iteration

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.166147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.574853Z digest=sha256:44e787c0bfee49c1667d39b83031c1566315158adb2aef2407439beac75cdfa8

Observation 1f620bd3-35e8-4b44-a8e0-c92594300be9 · outbound

This paper cites Semiparametric theory and missing data, volume 4.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Semiparametric theory and missing data, volume 4

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.578347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.578347Z digest=sha256:740c08c196de8f30922711e21f4d37e2cf63dec25fe1c41273b433fc90176472

Observation 5253eba6-d85a-49eb-a3e9-63f1629d1595 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax weight and q-function learning for off-policy evaluation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.143639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.584284Z digest=sha256:c9829ace3a6631b0ae5ddeaa32471abc981c2cc0d6ca1c5144cd2c3bdae165e5

Observation fdef9586-60e1-47ab-aaa9-a545b885a559 · outbound

This paper cites Asymptotic statistics, volume 3.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Asymptotic statistics, volume 3

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.129716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.588943Z digest=sha256:989420c2476c826fa4f3e32424de658a26bebeaebce7ea8600c154e9e954d5fd

Observation f87c33c9-fd40-4ca5-8ba6-c586cc08db69 · outbound

This paper cites High-dimensional statistics: A non- asymptotic viewpoint, volume 48.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures High-dimensional statistics: A non- asymptotic viewpoint, volume 48

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.116170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.593223Z digest=sha256:94057b86c910bc4d3a42977847513fb1628ae1fb6e017caaa6beb445d7798c13

Observation 7b684ee2-762a-4cab-bb9b-37310552c199 · outbound

This paper cites Provably efficient reinforcement learning with linear function approxi- mation under adaptivity constraints.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Provably efficient reinforcement learning with linear function approxi- mation under adaptivity constraints

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.102768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.598331Z digest=sha256:b1162cd76510776a99f1c79462c2d260b85bb533d7726de421ec6530f2ff9dbe

Observation d95b0957-9ce6-4423-818d-b7197348b8e4 · outbound

This paper cites On gap-dependent bounds for offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On gap-dependent bounds for offline reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.088796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.602873Z digest=sha256:d1068444497df4339255f301788fef7872185ac026d24ce3a481d0b7387aca33

Observation 0646dc31-11d4-48ce-b901-1af3c63df95b · outbound

This paper cites Opti- mal and adaptive off-policy evaluation in contextual bandits.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Opti- mal and adaptive off-policy evaluation in contextual bandits

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.075476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.607201Z digest=sha256:3c12c4882beb693981f6817f0edb3ccb132e233beb25d351b2631b67b9679fbd

Observation 17268b18-fa42-4ba1-9a94-47bdce98e4c1 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Behavior Regularized Offline Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.611265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.611265Z digest=sha256:8b82a748713d18eae0e4e5fb7c68407973d88b983d039b15c62569a5b611b8ca

Observation 4c6f5143-4834-4b7f-b907-f772f2121fc5 · outbound

This paper cites On the optimality of batch policy optimization algorithms.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On the optimality of batch policy optimization algorithms

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.615198Z digest=sha256:92e61e591fd121b4e16af462d3ec35edfdda57418783f47b58938a2cec5eec85

Observation 9757f4e0-0dbe-4062-bba3-b3cfcc744b21 · outbound

This paper cites Q* approximation schemes for batch reinforcement learning: A theoretical comparison.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Q* approximation schemes for batch reinforcement learning: A theoretical comparison

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.047894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.619015Z digest=sha256:fbe0664893ee030889c8c5508d464c70d5f490489baf1eec7951a1a8674afe9f

Observation 189aff4c-fe89-4510-a5a7-af8d297a211c · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginal- ized importance sampling.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Towards optimal off-policy evaluation for reinforcement learning with marginal- ized importance sampling

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.034006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.622611Z digest=sha256:7a7b67b7fdc351dd1a7b6812709bf33f5dfa8f6b6c3012a656e43a9708b7fa1a

Observation da2895fe-2f71-49aa-bfa3-6016f79f0d28 · outbound

This paper cites Bellman-consistent pessimism for offline rein- forcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bellman-consistent pessimism for offline rein- forcement learning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.020810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.626325Z digest=sha256:f3e608552f7149ea9eeb3718133186dd215a5fe14927c1a2c693ea288650271d

Observation 19817067-7e7a-4a91-9f41-05d340030dd3 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.630025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.630025Z digest=sha256:711be6b5c178f7940cb9681a52bedff83a85be0202f0d4385223eb31009a14b7

Observation 6064c2ed-6b28-489a-8a2b-74a6bfba6e3b · outbound

This paper cites Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.633815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.633815Z digest=sha256:06a91845cb2ddd3acbc517f30827d566836674c3556b2f4fa3b9a17438612f44

Observation 6dfea6f6-3eb1-4951-bbdd-83da7727c7d5 · outbound

This paper cites Asymptotically efficient off- policy evaluation for tabular reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Asymptotically efficient off- policy evaluation for tabular reinforcement learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.998146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.637934Z digest=sha256:02db3bae5279b8e22036d309bbbada62aca35f2a9dddf0ffd5bd32950d282882

Observation 80472f83-d677-47e9-aeff-1e7a3590bdcf · outbound

This paper cites Towards instance-optimal of- fline reinforcement learning with pessimism.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Towards instance-optimal of- fline reinforcement learning with pessimism

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.984804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.643262Z digest=sha256:483adc50680bee72fa8fa049e271f6efbd621f76376f2188f72874ff8abd9f89

Observation cc4bef7c-c292-402e-8207-a764985f2dcf · outbound

This paper cites Near-optimal prov- able uniform convergence in offline policy evaluation for rein- forcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal prov- able uniform convergence in offline policy evaluation for rein- forcement learning

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.970026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.647430Z digest=sha256:a2f7772c6f77643543a84561a319517d2f2e70e85683f7541df5a0e626e00a47

Observation 5397d07a-78d9-4e51-baa1-2aa64fa26568 · outbound

This paper cites Near-optimal of- fline reinforcement learning via double variance reduction.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal of- fline reinforcement learning via double variance reduction

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.955788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T22:19:16.651877Z digest=sha256:f7ec1a97cfc392893e96180779a993b9c0d5e28e98a305e8335768eb51fb932e

Pith citing papers

No inbound Pith citation observations are available.