Pith. sign in

Paper Citation Record · LEDGER

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

As of 19 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2606.31769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31769 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:55:27.313461Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy84
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14d4c429-7d32-4c9f-8a3e-03f50fc6269c · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.239431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8f594c630c7a37fce44c02df3c55fa8d4326fcc8abc06d5a67b49378f98d046e

Observation 73453eab-37e1-4f8b-967b-65daaab3520b · outbound

This paper cites Conference on Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Conference on Learning Theory , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.381207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2428d25ff25c9aa53a7397049103155fb5538855ff0d8246ecf84b1ba670585b

Observation 5db8998b-34fe-4760-8bcb-c3c0c5d4d7c8 · outbound

This paper cites Near-optimal Regret Using Policy Optimization in Online.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal Regret Using Policy Optimization in Online

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.341134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:542e94ff659da623bfe4091931c169ef07cfd7fa39539a67f7977876d1eff947

Observation 5c621f20-e926-409d-ac62-4cb6a505d1c8 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.345951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8951fd67c0029b3f906b913f8ff339008db5c485e80837d5afbfb80174a2208a

Observation 05a895cb-c2dc-4d63-b849-3a6d1a0eb3f3 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.339102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2cb4c7325ca30ec185a4c71a380fd47ead265b31b3ca6fff4088678435b69d5e

Observation 20323228-acb9-4341-9b07-f2fe5683d9ed · outbound

This paper cites Proceedings of the 34th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 34th International Conference on Machine Learning , pages =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.334006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:51f0200591863df2df3b2a24aacc8d6224acedd1ccd9f3bf59e0a4fc34709fd2

Observation b571c96a-376b-4eaf-bbdc-16a809693033 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.269768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:28bc54c4d5d34b510ff49ecc76ae97edf989154511897d0e7080e086194b03d2

Observation 85d2836e-8a92-4c80-b93e-fb39356e84a9 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.326980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1036b4215001c6af7660e5e8aecf4179538280ff72dfa181f136244a10cd04c9

Observation 3225afe0-49ef-4e5b-8a53-0a08dec03ee8 · outbound

This paper cites Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.397181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:aa78cec3a40be4df16fc59e60987de99938bfc5eae72a9f90539834915eb3f63

Observation 9ccd5974-909c-460a-b790-a20d1b0c7391 · outbound

This paper cites Non-Asymptotic Gap-Dependent Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Non-Asymptotic Gap-Dependent Regret Bounds for Tabular

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.361225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:cb7c206ab3cccb5735a46b56188af07ece85d48ef7786a5ecbfc023348162c5f

Observation b960668c-5eb9-44c5-b74c-b90e2bfd2736 · outbound

This paper cites Online learning in episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online learning in episodic

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.359031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2cba0f8dcbd13cdd009b4264187fe3111922ac3b10378aee07248f7255f7314e

Observation 5a1fb6b0-2645-4226-a0fc-8d077f6a329c · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.319840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e31f2dd9bebb9e26de1379975c57019a1202d77e8ece597bc2e8877be92d5575

Observation 45594917-e06d-48fe-81ec-bc38916cfaed · outbound

This paper cites Learning Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning Adversarial

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.268635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:bf62e8aa328e38a689a1f715374c731c3954de7d8dd024beb95eceeffaaa5cb1

Observation 1757f675-696d-435d-adb6-25d277b97c66 · outbound

This paper cites Policy Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Policy Optimization in Adversarial

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.316384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6dbb8dcef9128229f79137e35cc1d7e30b0525b97743e47466a0ed47c8e302cd

Observation f194eb67-45c9-4105-b804-02d92d4f7fab · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.370981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:62695d31ef5a09ab56d18620987881967f264c4cb087cfd15797c1980a4167d6

Observation 6eb8e578-5ca1-4ca1-948b-67c146c569f9 · outbound

This paper cites Laurent and P.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Laurent and P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.312811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:037c21dd99434cb92645fd1cb017be523fae4cec1e2af83ac3e3088aa0968a74

Observation c0992d9e-a455-4916-839e-cfc95685db02 · outbound

This paper cites Refined Lower Bounds for Adversarial Bandits , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Refined Lower Bounds for Adversarial Bandits , volume =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.259880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f49e7bda88511c3bf4f7e6c73295372c2edaa1615a26235c987b4f6f03dcf2d6

Observation 0bdcf0ce-552d-4d47-8257-9bae3ad77771 · outbound

This paper cites Probability and Mathematical Statistics , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Probability and Mathematical Statistics , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.321775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:99580749ab22683b48273731829d85958e43159be1f223d97a5fd4b94ec9c231

Observation 33c6b8d8-f7ab-4633-9ba0-64cd2dbc7cd0 · outbound

This paper cites 2013 , month =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2013 , month =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:833b1910ce0cc2d89fc8c60c3c9f4c5f090a698c8e83279206a3d38ed5377b29

Observation 1fa971e0-eb76-43ac-98ed-64bf069def4f · outbound

This paper cites Tuning Bandit Algorithms in Stochastic Environments.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tuning Bandit Algorithms in Stochastic Environments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.332323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:502bdd3abdcce8dc3ba77d4ed1a84c9e7c56b1ba30d4aad331fedf50773a96cf

Observation dad72bcb-cd68-4dcc-a782-f0aab952156a · outbound

This paper cites Online Convex Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Convex Optimization in Adversarial

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:25336a458526c5c05d3568b50f459e960b3852c72d93fcdf77e0fa5076ccb02a

Observation eec37106-7870-4a44-9e85-da8a6caff019 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 37th International Conference on Machine Learning , pages =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.292144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3d9c48ae659f076a3d180f5b84c345343938f79e21644509221961bacca77c6c

Observation d7911cad-e059-45e1-a737-dd470191f0d3 · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.223530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:bef166ba86e6f766b6c807f42a927f78c2e81e0a76aa97d446729354fdf91a6a

Observation 18fba09c-b30e-4ca9-9633-30eff050de10 · outbound

This paper cites Online Markov Decision Processes under Bandit Feedback , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Markov Decision Processes under Bandit Feedback , volume =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d83d188e6e7a88f712b58d2633656e2d4cf883e1df2d1252104f79061b7d1255

Observation 3c4833ef-90c9-4662-99dc-7722f431363e · outbound

This paper cites Learning adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning adversarial

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.403917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:dcab306f458b4694052d78a79adb64807cd689f4ce5b8a85a3e2d125e75ef179

Observation 8aeca06b-870a-4d02-9d44-e1de777d92ef · outbound

This paper cites Near-optimal regret for adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal regret for adversarial

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.297517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f39c0fd0e4d8dab8abfc1768311367565b17c0e7e8c1518c9fec817fa4f50975

Observation ef4a2f62-5098-4b17-9e9a-c3301233296b · outbound

This paper cites Delay-Adapted Policy Optimization and Improved Regret for Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Delay-Adapted Policy Optimization and Improved Regret for Adversarial

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.275251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4d64f3f54288335910ed8f69f7bdb41bb349adf03f12e645d16ca57203fd7feb

Observation 0e56bc8f-e625-4ed1-88c8-0fa5f1e6c559 · outbound

This paper cites Finite-Time Analysis of the Multiarmed Bandit Problem , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Finite-Time Analysis of the Multiarmed Bandit Problem , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.407915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:22cace35b21330c1db31711a27d6c9eb99e9cbfcb453b1ace4e3bb0352f68afb

Observation 55f4fc67-b198-4ae1-bcde-4649e38d84c1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.401904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:40e63daed93f93c949d7b857bdebef0045efb719531164bc222b4aad52001f5f

Observation 79b11b62-41de-4c04-958e-5faa2952870a · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.301336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3a510e1468fb2b8b20dd15598de73edd4aadc443106ac0e52f8be9c8e40754be

Observation 61c2ad1d-8f56-4911-8464-6682442becd8 · outbound

This paper cites Simultaneously Learning Stochastic and Adversarial Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Simultaneously Learning Stochastic and Adversarial Episodic

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.314549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8d5c1333990c4fef6b09cd9a1efc56f47f3595fd7cfc3ce58764a80d6ac8cbf3

Observation ead8046c-1c24-467a-bf8d-e9d780a30add · outbound

This paper cites The best of both worlds: stochastic and adversarial episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions The best of both worlds: stochastic and adversarial episodic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:38702292d39d0f86c8e62f530bc2bd3ec43a6bda4b527a8e771769a2d428247d

Observation b7a99ceb-326d-47e8-8956-00025523569a · outbound

This paper cites Proceedings of the 31st Conference On Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st Conference On Learning Theory , pages =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.273202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:fa910738fb93e6aa9309ea5093a3b87a9a90ff5066cc5fbdb8af7b02f718a379

Observation 7b6acfcb-4b3a-458a-82c3-a25c9792852a · outbound

This paper cites Tsallis-.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tsallis-

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:918c8a0b6d52a6cd1dbf12425e9376c2b5f78a18e46446177063da9c6b7c6697

Observation 283ef0d8-b7d3-4af1-88de-a7227d4cc57e · outbound

This paper cites Improved Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Analysis of the

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d1c170dd3398721ff220101185dad132382cfa4c579f911f5b752dcb3aee441d

Observation 1ad86420-6132-4b63-bdcf-211a82a68f68 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 36th International Conference on Machine Learning , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.335767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6e0286f2a6ebef71cde06cab4d9d3b5c0513f6e4c8c150a39d2306529517b9f1

Observation 069f8c57-4db0-4a62-9e12-88cdce021f82 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.385429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8ba1a1d759c389be3dfbc65db6fdb109040517f4085b57d1adedca0401ae51c5

Observation f3a43589-aff7-4c0a-9851-de826d671e9d · outbound

This paper cites Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.309320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3a36b0a061d85d2949c2799018275bdc610a958e1b932f231836e30003088987

Observation 65157022-e1de-4bf0-9fcc-b32a67f06990 · outbound

This paper cites Proceedings of The 27th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 27th Conference on Learning Theory , pages =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.289311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2e9c5e75f5d7805c8b843cb9baf988f521d64865498f5884cf34d47e0ae21529

Observation d29675b0-ca26-420e-8511-4cd3846ba2eb · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.230166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c0145be8f3421054ddfabaf79ba9c8962984ff92b80683ab592584f716dc2319

Observation b80e341a-a2d8-42a5-9890-bf1c7145cf5b · outbound

This paper cites 29th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 29th Annual Conference on Learning Theory , pages =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.295280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7c36e955524258b1b04d5b27badd2045bbdc5266434a64483fd887a95332efbc

Observation 88bc35d7-fde4-467b-aa22-8d35e6dd8b77 · outbound

This paper cites An Improved Parametrization and Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions An Improved Parametrization and Analysis of the

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.237992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:999ef63b014a8aea49a33f17a339f3d451874d61dd66c68d4073517245205654

Observation 1fc76744-10bb-4bad-9298-0f5893c21a12 · outbound

This paper cites Adapting to Stochastic and Adversarial Losses in Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Adapting to Stochastic and Adversarial Losses in Episodic

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.266508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6871d3807fa2f5dab775f39443a31338cff2c5559e307668a5630dd5b39b97dc

Observation 30f6b5b5-eda3-4ef6-a4b1-46800372c064 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.373563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ebc1ca1ea052c1deced978110fd6b2c0080dc9bd7ee68f2ba7d555b30436ba12

Observation ce6b0064-2ed7-49a0-b78b-04f67489f4c6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.337481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1bc84a307cb3f65d69de73fb0240494cb3b5f85802922a5840e45d86222a6d30

Observation ef46ec24-05b7-4742-ba9a-e616892f0806 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.402086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:888a1c9d3f0667e5130781490c5baa629b60cdc1222031421df84568ae3ee803

Observation c8b63570-73c8-4d00-9521-052f06db5f20 · outbound

This paper cites Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:35d3858f848927392fb5554a0a3a7a5687e2368bc2aec6c78aa54130eb93f822

Observation f761abec-1550-4e4b-8410-5c0bfe4d2265 · outbound

This paper cites Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.892402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:09f4c81064120c19f3568d524bebea82e4ed1f011f45d397fa79108963403d92

Observation b079470e-a09c-47a0-bd46-652dc5cec595 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Artificial Intelligence and Statistics , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.228386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:335f6181a19a03fd0b2bca453648951c72148b23ecbba29937c370d9fabfc5da

Observation 7351568b-25a9-4068-b13f-e3efda3a9fad · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 38th International Conference on Machine Learning , pages =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:cd348c70d7e2fecb683b2ad6e0b04563f1e49ff9d519e574693769657f4eb83e

Observation 975437ad-7c4b-4ea9-8ca0-5cc0b847ecbb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:be17aa195cc95f401deb0d24772adab8cfd74093c16287ce05a22cab9a33e86d

Observation ebe4fd21-114e-47a6-9268-908d3f87c90c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.323736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c663e21f76b96e350acd623f4dacbe3f0713f70fe03f59b297cde77693c9be8e

Observation 18ef48be-ef84-4f11-ac6e-5a0301de82d1 · outbound

This paper cites International Conference on Algorithmic Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Algorithmic Learning Theory , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:20745d43824577a3daa4ab6ace839c69668a289a9047202f3693692df1c220a4

Observation f5336e48-97d2-452f-b844-47925bb868f2 · outbound

This paper cites Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.271271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:343f7d735f18f710f5cd036b726a1fd7714c737087701b6f43f481839447d376

Observation 49eaa7ff-e90a-4c68-8843-76c717c97ab1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.248384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:72d29db12b59ea331a35b0e985992051bb30919da3403cb7ef3df8236ae54921

Observation 08eb070e-a267-4f22-a582-9e9212c0c8cb · outbound

This paper cites Proceedings of Algorithmic Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Algorithmic Learning Theory , pages =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.299423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:efb671ae586abaeccaa391fcd4461da9457e069f44ad3522b9f6677f2db60a92

Observation 0b9b304d-242e-44b7-9f9e-5a9ae06d0959 · outbound

This paper cites Proceedings of the Thirty-Second Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Thirty-Second Conference on Learning Theory , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.388785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d49413761df2bbb59544ed1511b06060fc8a69b10f7ec453bebbcdc094d2b7cb

Observation b2b371ee-f2dc-41f4-a3ae-5e9bcb46a8b2 · outbound

This paper cites Proceedings of the 28th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 28th Conference on Learning Theory , pages =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.390680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:14a8e7203abaf29ca525114a46e2f5d669e07a43950d0beceb0db804056f81a3

Observation 2719bc76-2c44-4fdf-8959-5355e859fa57 · outbound

This paper cites Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.277184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8c798daf962990c353495cfb63fcabe8b93001445aeb6da9df453c8dc06934fb

Observation ad6ab9f6-3957-433f-bbf8-240cbc4db948 · outbound

This paper cites Gap-Dependent Bounds for.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Gap-Dependent Bounds for

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.330346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a4f1c7585f4bc8585d4d7ee466f1d68e74f5362e741b07cb1b93d27f2621d0a9

Observation 734b30cd-ee00-4ded-b903-214fa1599dbd · outbound

This paper cites Proceedings of the 26th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 26th Annual Conference on Learning Theory , pages =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.258818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:05b098d59a90adee9c93dee5d3b77de13e58cedbb88741435ef946cb75051959

Observation 7a008eb7-8a0f-433c-aab6-24c3fe8865e9 · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.392809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1ebf911e657cb7680263b3c5d53201b96ae7d0c3611ce5554e904a19e916a5bb

Observation 05a66b6a-d885-4d27-bf79-232c002494e2 · outbound

This paper cites SIAM Journal on Computing , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions SIAM Journal on Computing , year =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.264201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:69a5e8403625eda2c28c5da9516476c875742ec5c94db880a9c1d9c61da9b4a9

Observation 16a038d2-a79f-469d-afdb-a9a93c832959 · outbound

This paper cites Proceedings of Thirty Eighth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Eighth Conference on Learning Theory , pages =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f090bfa8b48e9c0423bfa77d7db923e739a2dd0abd5c024eb4af0b4e1264686d

Observation 1f334187-8f02-4e90-b1f0-4194950b9dd9 · outbound

This paper cites Proceedings of the Nineteenth International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Nineteenth International Conference on Machine Learning , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.291102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8634a467df68362f1a09f85493b1e9720deca568f947981f01ab532c3b7edb3a

Observation d1c1ef9b-7c8e-4ab8-87f0-75b1bede660e · outbound

This paper cites Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.386464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:44477f1a6812bf7424cc1d921162524b29d60adfc2728eb66e7ca4906bf41eac

Observation 46308e1a-e16c-4827-8310-11d5b19ee355 · outbound

This paper cites , author=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions , author=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b19b28a27df619fab4d0133d10409b6d0c5c2f961140d424b7c404e426b6354a

Observation b9eb6bdc-d7d5-4fa2-9b68-cd3649abfeab · outbound

This paper cites IEEE Transactions on Neural Networks , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions IEEE Transactions on Neural Networks , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:27ce04c3d580af4b3c6389bfa23fa1ab1711f4e2370920003a5ff210b2c73055

Observation 56f0e67a-ed8d-45aa-aa73-4392dea58ee7 · outbound

This paper cites 2003 , booktitle =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2003 , booktitle =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.404206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:10a10beb0ee022807a1b418f96a38e6c2f9beed60c885055347ad2fb69578ac6

Observation c2186f21-3fad-4ae9-a1c9-031e754cf68f · outbound

This paper cites Information and Computation , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Information and Computation , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:36a5b3705b8244f59fb943912eca1a1ada6ca8fdf8078b7c20c2e8df64c9e6d6

Observation d0a6b86b-a741-4147-bdca-1198d3ff7a38 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proximal Policy Optimization Algorithms

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:55:35.889593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:27f42e19eec74721726f0d2fac57a7219d85ae0d68418b123e68430673fa5d3b

Observation 5dac5bd2-ce00-4818-83e5-6edfbfc22023 · outbound

This paper cites and Veness, Joel and Bellemare, Marc G.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions and Veness, Joel and Bellemare, Marc G

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.256923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:414cad9d3fd6a1c62c8ff3fc6c95078bb3ae7abe2dcb8f9887f27e4682a726a7

Observation 64ef28e4-d109-4630-8e76-61f152b6f25f · outbound

This paper cites Nature Medicine , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Nature Medicine , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.372950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c868e93f9667fd4e0293176d250d5167f76926e2c2fa8c6919b1ba3cb2500951

Observation 6b2b7331-052d-422e-913b-d221591b1cf7 · outbound

This paper cites Empirical.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Empirical

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.328708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:808d764bdb5a0664ad8af87ca467c33c2848de4a3a0f78462f1503393a0a552d

Observation 02221255-3002-405d-9dea-7106dafe63e7 · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e1e1d204c8a95e3aa93285d5e46355d9bfa048ccce379ffb4b67bf8cae8383ba

Observation 0e3c5536-7364-4b30-bac2-00b4f2c0ab60 · outbound

This paper cites Reinforcement Learning from Adversarial Preferences in Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Reinforcement Learning from Adversarial Preferences in Tabular

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.279149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4c6a79d9ad2fc6736d6aeaa9e2c35d64a8df24e674c5795190e95dd5a05f0e74

Observation 048475ad-6b6f-4e6d-82b6-f8446350cf73 · outbound

This paper cites Proceedings of Thirty Seventh Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Seventh Conference on Learning Theory , pages =

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.288156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:53d0610c6a4f09b907d4282cb93f60de56f6cca4b33c0dece6753e251918b84d

Observation 328d5ee9-e3a3-488f-a8d2-9aea70f402ed · outbound

This paper cites M arkov Decision Processes with Arbitrary Reward Processes.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions M arkov Decision Processes with Arbitrary Reward Processes

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.307301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1b4c3ecf89d4a922f809d61beac90e92b21b0b5bc8b6413a77020f5166d7c9a2

Observation 07fd084d-5d60-4008-855a-81960e3a7ce2 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.311024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:83b17e28b7305d73b8293f515481a9e1af278e74550a362736dce6b1dbd5d571

Observation a85d3468-c1fd-4e7b-b48d-8c1283dbe9b1 · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 39th International Conference on Machine Learning , pages =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.234774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ae3ab72ca7ff83aebe6a50a7593d63248ce47f0ff2a0d630a043cdf4e85443df

Observation e9096485-519a-4d85-8e7e-de228c345a8e · outbound

This paper cites 2009 , author =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2009 , author =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.305606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0a99fa4135b06e9839382b5acde7f9c4fa65e680a831d96cb3670d1ec2f2abcc

Observation b800689a-e08e-43bc-b1d7-7f7352943bba · outbound

This paper cites Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.363486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7a788b5e8972e41b9efab9d59d3ce8bf52f3ac0f5d767d5922d1c176b7159014

Observation abb9c9b2-1291-4586-9c3f-a45891fbcc19 · outbound

This paper cites Episodic Reinforcement Learning in Finite.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Episodic Reinforcement Learning in Finite

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.355915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:055afab50664106255cfd602fadff9dd0cedf2b14385a5d254fbaaa682a3be6d

Observation 2604cc15-1e64-45f0-b27d-a2c119553618 · outbound

This paper cites Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:bd0b21f81bf9ed74d951eb12579c03f5cdd1847fdca8494e0ca473f2da29379e

Observation 6865871f-6166-4178-ad00-ffc61fb2db14 · outbound

This paper cites International Conference on Machine Learning , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , year =

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.354755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:74f63eb2d80678eec4e57a15751872ad6be1222e0aa951dc0eb845c0671ca533

Observation 5d6dfc3b-f803-42a3-a658-c37bc2f2bb6f · outbound

This paper cites Narrowing the Gap between Adversarial and Stochastic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Narrowing the Gap between Adversarial and Stochastic

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0ffc4bfc69374ce3ef0c417f4e1cb615c7aedb8e8a8d390ec6159f9005099272

Observation 40ba6c59-3631-4677-ba57-2d8ab9a7cc61 · outbound

This paper cites Proceedings of the 32nd International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 32nd International Conference on Machine Learning , pages =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.318135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c40cd54ae054cd9538b87caaded32243ab744187cd02f6b6607bf074ff6eff59

Observation 1283848a-a0b0-449f-91e3-da4a95ad938c · outbound

This paper cites Advances in Neural Information Processing Systems , publisher =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , publisher =

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.343968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:79e1fe7e73ad3a4922be0cce58e086e975c2c743245ae01b179b80272c63a644

Observation 65a0f318-5b1e-4248-81bb-e8d6426791b9 · outbound

This paper cites Fine-Grained Gap-Dependent Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Fine-Grained Gap-Dependent Bounds for Tabular

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0d56522f9c3d3b98f980641056b3ed9013dfb5bb2bee1fb840a47d6ad6b3edd3

Pith citing papers

No inbound Pith citation observations are available.