Pith. sign in

Paper Citation Record · LEDGER

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

As of 19 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2606.31769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31769 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:55:27.313461Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy84
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14d4c429-7d32-4c9f-8a3e-03f50fc6269c · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.239431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6e2dabdd46bd79c4ad5c65e94763da437d2b2a59d00d64464c06305c09b2e156

Observation 73453eab-37e1-4f8b-967b-65daaab3520b · outbound

This paper cites Conference on Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Conference on Learning Theory , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.381207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:560b7d90b7f2fb32c0d50f9c205422d69f12ceaa5fd82fa3b50b16830dfcc96b

Observation 5db8998b-34fe-4760-8bcb-c3c0c5d4d7c8 · outbound

This paper cites Near-optimal Regret Using Policy Optimization in Online.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal Regret Using Policy Optimization in Online

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.341134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a818dac14ee62df015e24db14005bf570c7a275030e2348741d3ed3d465c073f

Observation 5c621f20-e926-409d-ac62-4cb6a505d1c8 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.345951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:680c0ebe2a9df66eb821a02a349117b0dba71bfc309af9e2c4ebe428314867cb

Observation 05a895cb-c2dc-4d63-b849-3a6d1a0eb3f3 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.339102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:17bdef0b956497f38387c7469854fac4c83f0492b57ee421f826fc09a89358e4

Observation 20323228-acb9-4341-9b07-f2fe5683d9ed · outbound

This paper cites Proceedings of the 34th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 34th International Conference on Machine Learning , pages =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.334006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9390fc8103cd39a139a8e0d8395be3b9a4b4270a24a12b1d49b2cdafebccdb6d

Observation b571c96a-376b-4eaf-bbdc-16a809693033 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.269768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e7490e131e858fce595babe9398ade34d7ee24fd1040294dc971a4d44825e954

Observation 85d2836e-8a92-4c80-b93e-fb39356e84a9 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.326980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f118dd52cb58f5163a5604b3f345c30f1c24d6607e5073befc02843d7752d5ea

Observation 3225afe0-49ef-4e5b-8a53-0a08dec03ee8 · outbound

This paper cites Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.397181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:dec06defd5497fa54b02662d780c33e76fb8cfa199c0f412cae6397c526bcb0a

Observation 9ccd5974-909c-460a-b790-a20d1b0c7391 · outbound

This paper cites Non-Asymptotic Gap-Dependent Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Non-Asymptotic Gap-Dependent Regret Bounds for Tabular

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.361225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:95a93728c849daa99d6fd792fa8356c68883ba09c686c4095dd0eba68bb75219

Observation b960668c-5eb9-44c5-b74c-b90e2bfd2736 · outbound

This paper cites Online learning in episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online learning in episodic

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.359031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d047f6265d4d2539207f1417a75c26cd9b168cb91cfa788f1485ef21f93af4ef

Observation 5a1fb6b0-2645-4226-a0fc-8d077f6a329c · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.319840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1e3e3b9ca1aa5d2d91c09b5e37c4d3b7305e9cd236106fb581a8fd106dafe192

Observation 45594917-e06d-48fe-81ec-bc38916cfaed · outbound

This paper cites Learning Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning Adversarial

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.268635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3fa1d007308c3e62cce3ffcb7341521213577b507080ac4a0084bb815d67e125

Observation 1757f675-696d-435d-adb6-25d277b97c66 · outbound

This paper cites Policy Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Policy Optimization in Adversarial

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.316384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:585819205fc2fb45910a81c34b0a1a6732e9928a87c201e44fd6257afb07ec08

Observation f194eb67-45c9-4105-b804-02d92d4f7fab · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.370981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:32e4d3954c09add0077ff583d1b56a6f9e6d551cd227ae1d33b7b20fd4d8eef7

Observation 6eb8e578-5ca1-4ca1-948b-67c146c569f9 · outbound

This paper cites Laurent and P.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Laurent and P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.312811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7ec3b89afe3ab18ccf000397ab21428ccfaf70264c6cad82e38570800c8e869a

Observation c0992d9e-a455-4916-839e-cfc95685db02 · outbound

This paper cites Refined Lower Bounds for Adversarial Bandits , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Refined Lower Bounds for Adversarial Bandits , volume =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.259880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8cf0246f8f4b6bbdd89ecc4b9c819f656a35431475b6d83450891ed95020b485

Observation 0bdcf0ce-552d-4d47-8257-9bae3ad77771 · outbound

This paper cites Probability and Mathematical Statistics , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Probability and Mathematical Statistics , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.321775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:df23b78ce3c933dac7cc5d770a4c5efe55c3622cc202cabefa12cd055a1bd499

Observation 33c6b8d8-f7ab-4633-9ba0-64cd2dbc7cd0 · outbound

This paper cites 2013 , month =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2013 , month =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9867fe514141ba59fc3f3235f1940d6c2d7b6b0b9a1eb5d6786ffedcc800cfe7

Observation 1fa971e0-eb76-43ac-98ed-64bf069def4f · outbound

This paper cites Tuning Bandit Algorithms in Stochastic Environments.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tuning Bandit Algorithms in Stochastic Environments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.332323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1f9c30f4149863aa650643876e45acded5561b8195f85e9e2302a5c41b62ddb2

Observation dad72bcb-cd68-4dcc-a782-f0aab952156a · outbound

This paper cites Online Convex Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Convex Optimization in Adversarial

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:01c417831cf7af06d48d8adac9f93535d7aed89c43a0c0fd71b6f963775b3dd3

Observation eec37106-7870-4a44-9e85-da8a6caff019 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 37th International Conference on Machine Learning , pages =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.292144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e82635ee2c0403e4b26808920c53287f12d11a5ab14c29dfe4df9d55c1b86a1e

Observation d7911cad-e059-45e1-a737-dd470191f0d3 · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.223530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:76290f4e6859e8e31c738b33c51c7764baa415e2efe3a841a7fc93fc42cbb902

Observation 18fba09c-b30e-4ca9-9633-30eff050de10 · outbound

This paper cites Online Markov Decision Processes under Bandit Feedback , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Markov Decision Processes under Bandit Feedback , volume =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f0492b5370af0099ca56652c7067cd034576be198b9a24193364cd9c36e2b440

Observation 3c4833ef-90c9-4662-99dc-7722f431363e · outbound

This paper cites Learning adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning adversarial

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.403917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c1eea86b65fcedd5627929eed21f1fec55174620dfb536fb6b6a301eeb15944d

Observation 8aeca06b-870a-4d02-9d44-e1de777d92ef · outbound

This paper cites Near-optimal regret for adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal regret for adversarial

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.297517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c0c87995b5b609920050567fdafa42aa204eabe136901fad769ebb15184352c2

Observation ef4a2f62-5098-4b17-9e9a-c3301233296b · outbound

This paper cites Delay-Adapted Policy Optimization and Improved Regret for Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Delay-Adapted Policy Optimization and Improved Regret for Adversarial

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.275251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ccb5d3444dd8518e1a722163eb8827a56d627b7ccbbf3b9682aacd159d75d96c

Observation 0e56bc8f-e625-4ed1-88c8-0fa5f1e6c559 · outbound

This paper cites Finite-Time Analysis of the Multiarmed Bandit Problem , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Finite-Time Analysis of the Multiarmed Bandit Problem , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.407915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:440708ff03ae6a49811d1a3e5bf969526fd908d7c7b1a121ef4e272fd970cc90

Observation 55f4fc67-b198-4ae1-bcde-4649e38d84c1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.401904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:46da2ff21ddb3824319b6fd0585fc4825a5a544efc216c697f5c017cbe198df8

Observation 79b11b62-41de-4c04-958e-5faa2952870a · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.301336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b13a965d3585ae2e6c879340ba7bcad8d65a732e9521077a94566d284fbc2c0e

Observation 61c2ad1d-8f56-4911-8464-6682442becd8 · outbound

This paper cites Simultaneously Learning Stochastic and Adversarial Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Simultaneously Learning Stochastic and Adversarial Episodic

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.314549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:20b8cd8c358b6045da8ef466617feadc389a822354bfe988b277143ca7b5d23a

Observation ead8046c-1c24-467a-bf8d-e9d780a30add · outbound

This paper cites The best of both worlds: stochastic and adversarial episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions The best of both worlds: stochastic and adversarial episodic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a9782f429e0068163cc162bed995357f5f51ce8a4218e03536dcc7c940a13438

Observation b7a99ceb-326d-47e8-8956-00025523569a · outbound

This paper cites Proceedings of the 31st Conference On Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st Conference On Learning Theory , pages =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.273202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:faff3acd0e3ae558abc6b3a846c18bddbceb9db956ec654198f197b590b56381

Observation 7b6acfcb-4b3a-458a-82c3-a25c9792852a · outbound

This paper cites Tsallis-.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tsallis-

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ccafbc57be87f91eada86936c25efd785b7ab54b6cff76568d87f54d942df25b

Observation 283ef0d8-b7d3-4af1-88de-a7227d4cc57e · outbound

This paper cites Improved Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Analysis of the

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a22aeddcac381d9882276ea96ceb9531a2dc1c2c0ebe5bdc9a9875b717d12f41

Observation 1ad86420-6132-4b63-bdcf-211a82a68f68 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 36th International Conference on Machine Learning , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.335767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1716c2aa6c6e84abacba48c74391aeca04b180b4398f961e4f5306c32b72d09c

Observation 069f8c57-4db0-4a62-9e12-88cdce021f82 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.385429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:13207fa1e9ce242b303deef7e42d9c04d9eaa5ba294ffe969e162d9ad41b49c4

Observation f3a43589-aff7-4c0a-9851-de826d671e9d · outbound

This paper cites Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.309320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b76e95ada0300c8eace55792b49dff0ebdb087dd7b094d31fb4aa1a0ab66e1ce

Observation 65157022-e1de-4bf0-9fcc-b32a67f06990 · outbound

This paper cites Proceedings of The 27th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 27th Conference on Learning Theory , pages =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.289311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0dbf18e7fbcbdc93701de69e89e40f64d7e885b6095be47fb91d50dc0e649a16

Observation d29675b0-ca26-420e-8511-4cd3846ba2eb · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.230166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4e70637cb08e73e268c595303095bbaca16db37d446d6c80531c382012887703

Observation b80e341a-a2d8-42a5-9890-bf1c7145cf5b · outbound

This paper cites 29th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 29th Annual Conference on Learning Theory , pages =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.295280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f87312cfcf574f092a7e43de95924c9f1741cf19f25b06140633e1e3b542d77f

Observation 88bc35d7-fde4-467b-aa22-8d35e6dd8b77 · outbound

This paper cites An Improved Parametrization and Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions An Improved Parametrization and Analysis of the

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.237992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:959fb1339ce98d3df4913e79311a6be2dd370b3fdf493c54c65e858369eed466

Observation 1fc76744-10bb-4bad-9298-0f5893c21a12 · outbound

This paper cites Adapting to Stochastic and Adversarial Losses in Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Adapting to Stochastic and Adversarial Losses in Episodic

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.266508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e19ad8c84b2f40c7ef675ebe3706babd0faaaf84203a8f52a894463ae551a64a

Observation 30f6b5b5-eda3-4ef6-a4b1-46800372c064 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.373563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a9531b398f6f0c6386e84486ecfa0c942d363b348604c9768dd1ef07bdbec681

Observation ce6b0064-2ed7-49a0-b78b-04f67489f4c6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.337481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0c0f7b3f03feb1bc176795c3fb64a7382fd7660570d022933d34aa01523056ec

Observation ef46ec24-05b7-4742-ba9a-e616892f0806 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.402086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:49a8db292c3d0f2b64b63546639ce1e6abfdc9d5885482c566be68cbd551317b

Observation c8b63570-73c8-4d00-9521-052f06db5f20 · outbound

This paper cites Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:eb0497ce6ecc0c9b94844240c28ff911bc63c594b094ca45598025f5beaddd30

Observation f761abec-1550-4e4b-8410-5c0bfe4d2265 · outbound

This paper cites Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.892402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8810a83b8a252a87185ab0a5bebfc742e532f8980e3ea7807bb0ce6d380f4655

Observation b079470e-a09c-47a0-bd46-652dc5cec595 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Artificial Intelligence and Statistics , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.228386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c0cec9365917cfcbdb0be783a709c9deb3354542ec160ef9c6babc525aedff39

Observation 7351568b-25a9-4068-b13f-e3efda3a9fad · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 38th International Conference on Machine Learning , pages =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:faa0a79cdf777d35ce1e463c30091957408f771cc1e0fa2fd9190ee51f0f6906

Observation 975437ad-7c4b-4ea9-8ca0-5cc0b847ecbb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:00e178839722357accecb18c274e43a2654fa8ba443cd0cc872bb3350ddd3272

Observation ebe4fd21-114e-47a6-9268-908d3f87c90c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.323736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:39a50a1f36274df3ff5a3c5b72b24c4d662df76abc000be76198e47bdb56cdbb

Observation 18ef48be-ef84-4f11-ac6e-5a0301de82d1 · outbound

This paper cites International Conference on Algorithmic Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Algorithmic Learning Theory , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c8ae706da8fabb5ac1beef52d91e24f7292fd09605018c89cbe3a33c30b99009

Observation f5336e48-97d2-452f-b844-47925bb868f2 · outbound

This paper cites Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.271271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:10df0023ec42b9a24fc53f7ea03680fb57b1604481cd284750b816db6080a069

Observation 49eaa7ff-e90a-4c68-8843-76c717c97ab1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.248384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a9c8a7443494e12c3a4f18658150fc62372e61e82eff6a9f455d8fdc09010fdf

Observation 08eb070e-a267-4f22-a582-9e9212c0c8cb · outbound

This paper cites Proceedings of Algorithmic Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Algorithmic Learning Theory , pages =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.299423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:82a9f023cce3ac4f26456da585f9e38df1c906c470f9904486893570b456fff5

Observation 0b9b304d-242e-44b7-9f9e-5a9ae06d0959 · outbound

This paper cites Proceedings of the Thirty-Second Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Thirty-Second Conference on Learning Theory , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.388785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:54b2bbfffede01ad795cdb774901e998d65118cc0eb5db6b5bcbe218fbf9632c

Observation b2b371ee-f2dc-41f4-a3ae-5e9bcb46a8b2 · outbound

This paper cites Proceedings of the 28th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 28th Conference on Learning Theory , pages =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.390680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c77c429c626c50a5b3b021238c08e582973dc9fcb8f33eab8ba9fea3c7511665

Observation 2719bc76-2c44-4fdf-8959-5355e859fa57 · outbound

This paper cites Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.277184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:51dbb5a041b6da534e86aa4dc0b6a609b732d273d24157555db3d0801bb6a55c

Observation ad6ab9f6-3957-433f-bbf8-240cbc4db948 · outbound

This paper cites Gap-Dependent Bounds for.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Gap-Dependent Bounds for

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.330346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0b253132d5fae41608d0708cac599cf1146b88449d8b76ab6824d572155879c7

Observation 734b30cd-ee00-4ded-b903-214fa1599dbd · outbound

This paper cites Proceedings of the 26th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 26th Annual Conference on Learning Theory , pages =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.258818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2fdea3526b9206a29de1047a09c560389747088bdc8ad66de0dc161455d204f3

Observation 7a008eb7-8a0f-433c-aab6-24c3fe8865e9 · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.392809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:666e27631dee9411808d7ca471b7a878930152c08dcc575dae26af2d2c0fa94e

Observation 05a66b6a-d885-4d27-bf79-232c002494e2 · outbound

This paper cites SIAM Journal on Computing , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions SIAM Journal on Computing , year =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.264201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:95995e7156b39d6a23cf9fac441b80729bc5f954d39a4a46eec158e6ba13507f

Observation 16a038d2-a79f-469d-afdb-a9a93c832959 · outbound

This paper cites Proceedings of Thirty Eighth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Eighth Conference on Learning Theory , pages =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:bf113587e9b451ad17af95ca3101314c81f431091035ac7360f7b510fde6ea0e

Observation 1f334187-8f02-4e90-b1f0-4194950b9dd9 · outbound

This paper cites Proceedings of the Nineteenth International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Nineteenth International Conference on Machine Learning , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.291102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d8d9d03d1317ecda13a5052a72cd97ac8372610c38ebea7c5568c4a3bc192f2b

Observation d1c1ef9b-7c8e-4ab8-87f0-75b1bede660e · outbound

This paper cites Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.386464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d17afccd33e91ff97d05a021233296c4c7a8f14a6d56a3fb30993d3c1361f6e2

Observation 46308e1a-e16c-4827-8310-11d5b19ee355 · outbound

This paper cites , author=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions , author=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b138dfa0668b23106268357128d414c2db2945f8c45021cb888232f7d1e63c2d

Observation b9eb6bdc-d7d5-4fa2-9b68-cd3649abfeab · outbound

This paper cites IEEE Transactions on Neural Networks , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions IEEE Transactions on Neural Networks , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1d569b435bf70a889d4459a7e119de2de06fd37bce67cf9f0eff673f677e22f1

Observation 56f0e67a-ed8d-45aa-aa73-4392dea58ee7 · outbound

This paper cites 2003 , booktitle =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2003 , booktitle =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.404206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6f5c07f939ddcea732e744c240ce93c12c7ea73556c2b5ade20f4561f4f1b0a7

Observation c2186f21-3fad-4ae9-a1c9-031e754cf68f · outbound

This paper cites Information and Computation , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Information and Computation , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:5742f1f23c25561619f17e4dcb7004a7716b212e6c62bf062703bddd397e8231

Observation d0a6b86b-a741-4147-bdca-1198d3ff7a38 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proximal Policy Optimization Algorithms

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:55:35.889593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8e8e75e093483553072d05191a94db7a5baab75ae3b44e4a2ffdc51b718fb385

Observation 5dac5bd2-ce00-4818-83e5-6edfbfc22023 · outbound

This paper cites and Veness, Joel and Bellemare, Marc G.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions and Veness, Joel and Bellemare, Marc G

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.256923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4702b88e3d8929154b617818e4a65ba035f937fb191819b68e99650c76f726d9

Observation 64ef28e4-d109-4630-8e76-61f152b6f25f · outbound

This paper cites Nature Medicine , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Nature Medicine , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.372950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:5a4dbbfaad85b387f0d567291826893082c44717f74bcfbc6d23fcd12da50e4b

Observation 6b2b7331-052d-422e-913b-d221591b1cf7 · outbound

This paper cites Empirical.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Empirical

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.328708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:dfbb7d5b40594b4769c182ee44cd84913dbe6b90463131a1a1b2d8426d1e04b7

Observation 02221255-3002-405d-9dea-7106dafe63e7 · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:18f7c7616e2f113c278c425a10dde086b212839cc7706d8f68fa5833973e90a0

Observation 0e3c5536-7364-4b30-bac2-00b4f2c0ab60 · outbound

This paper cites Reinforcement Learning from Adversarial Preferences in Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Reinforcement Learning from Adversarial Preferences in Tabular

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.279149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:bbd3975793d725823abd29806707369a3f982bb4161e7538317739e5479d9a7f

Observation 048475ad-6b6f-4e6d-82b6-f8446350cf73 · outbound

This paper cites Proceedings of Thirty Seventh Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Seventh Conference on Learning Theory , pages =

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.288156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ce1e5513ddaad24b6b3872b0b1a0a82dd3544b2dc82f09866bf37807d0fea004

Observation 328d5ee9-e3a3-488f-a8d2-9aea70f402ed · outbound

This paper cites M arkov Decision Processes with Arbitrary Reward Processes.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions M arkov Decision Processes with Arbitrary Reward Processes

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.307301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f1916a7a61011a46f1318e09d6258d7366934bbd69667ba19975a4ea0d74c2fc

Observation 07fd084d-5d60-4008-855a-81960e3a7ce2 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.311024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7acb973f212ea804ea76bd2deeb6ebea344bbf189591947e9cc520a0a833a47e

Observation a85d3468-c1fd-4e7b-b48d-8c1283dbe9b1 · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 39th International Conference on Machine Learning , pages =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.234774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:38de25eb86862cc9d4c9f04f4654811073cc440f3a80e0e31ce7119c8f6c65b0

Observation e9096485-519a-4d85-8e7e-de228c345a8e · outbound

This paper cites 2009 , author =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2009 , author =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.305606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8783b5720c4d2cbae720c21dbc43f024ceed1a5a00cf398a3868cc764c714bac

Observation b800689a-e08e-43bc-b1d7-7f7352943bba · outbound

This paper cites Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.363486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1a4b55dc2b6b8dc8e918fdbb0202008fc213521b34149a988e9861042208aeab

Observation abb9c9b2-1291-4586-9c3f-a45891fbcc19 · outbound

This paper cites Episodic Reinforcement Learning in Finite.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Episodic Reinforcement Learning in Finite

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.355915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2ef0a5b3275fb306cba69d860726d8f0ae8204d1094827aa0fd4a760af012f66

Observation 2604cc15-1e64-45f0-b27d-a2c119553618 · outbound

This paper cites Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ac3bd787895e0b1c976ed77ce716b0f2f9b75a61a03a39167e25e8e375ce7fec

Observation 6865871f-6166-4178-ad00-ffc61fb2db14 · outbound

This paper cites International Conference on Machine Learning , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , year =

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.354755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:590862d61374ce4fc516a97d30c07af458a2acb5532362d8f104ee1f60031cc7

Observation 5d6dfc3b-f803-42a3-a658-c37bc2f2bb6f · outbound

This paper cites Narrowing the Gap between Adversarial and Stochastic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Narrowing the Gap between Adversarial and Stochastic

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a49c60a0a77ad1fd83f23a450935d6ea1cecea4762b668e7c4ae735947c14a05

Observation 40ba6c59-3631-4677-ba57-2d8ab9a7cc61 · outbound

This paper cites Proceedings of the 32nd International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 32nd International Conference on Machine Learning , pages =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.318135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:db8d01bdb8203a511298a420010c2884b5165be0cf3c0f79866584850294670d

Observation 1283848a-a0b0-449f-91e3-da4a95ad938c · outbound

This paper cites Advances in Neural Information Processing Systems , publisher =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , publisher =

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.343968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:acd415373b284b9ff231376b21363f64f8ce1bb74413368321d7a3855c2b2fab

Observation 65a0f318-5b1e-4248-81bb-e8d6426791b9 · outbound

This paper cites Fine-Grained Gap-Dependent Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Fine-Grained Gap-Dependent Bounds for Tabular

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6e37419b9a0da69b8f7a2633a65f3a01b1fa78b1571b157361e0946e81a23ce4

Pith citing papers

No inbound Pith citation observations are available.