Pith. sign in

Paper Citation Record · LEDGER

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

As of 16 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2607.29294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29294 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:55:59.660672Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved83
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cd91f56-d330-429c-97ff-6865c814b3f1 · outbound

This paper cites 2018 , journal=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2018 , journal=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:50.602622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:50.602622Z digest=sha256:0070e596750172ab60a4e915c33c37705067e3c31f4379d12f8c2afb846899a4

Observation 67a4273d-4102-4062-86a2-b4650e51186a · outbound

This paper cites Advances in neural information processing systems , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Advances in neural information processing systems , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:50.654810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:50.654810Z digest=sha256:f5edc1ef96a9d6dc7ab351b2ca05bc59276e97f94b7ddadfe1902e903d6899b3

Observation 08b13f8c-0288-464c-b20e-2a56f2266dfe · outbound

This paper cites Proceedings of the 23rd international conference on Machine learning , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Proceedings of the 23rd international conference on Machine learning , pages=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:50.731457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:50.731457Z digest=sha256:5bb7c1394f5e6ca799d40c3ef4bbbbc129564ecf3af3104f8d0249997517951d

Observation b34a6d90-0986-4760-baec-2672dde9dfc2 · outbound

This paper cites Proceedings of the 34th International Conference on Machine Learning-Volume 70 , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Proceedings of the 34th International Conference on Machine Learning-Volume 70 , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:50.810636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:50.810636Z digest=sha256:4eb5bbbf61e5ee187711259f0898b5f4e93a7106518d54579fb65b1ecfb6c0a1

Observation 10200837-2709-4cc8-9c57-cd320248a92b · outbound

This paper cites The True Sample Complexity of Identifying Good Arms , booktitle =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification The True Sample Complexity of Identifying Good Arms , booktitle =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:50.884631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:50.884631Z digest=sha256:b34c3696df91c91b48897a6e040706da52c859fe929395190ad506f23af1ad03

Observation 3acdf1ad-28a7-4257-894a-7abc156278a5 · outbound

This paper cites 2019 , eprint=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2019 , eprint=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:50.950719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:50.950719Z digest=sha256:9f3c86d613fbbceec92e49d979f3fb780e1b0839c794be7c7dda6da54d3663ca

Observation d9e3a7cc-66bb-422d-b1b3-596bd31a89eb · outbound

This paper cites Conference on Learning Theory , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Conference on Learning Theory , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.028055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.028055Z digest=sha256:08f2c0251cf748ed30a01fd104adcd396907c417959626bdd656e098876e65fe

Observation a2c4b5df-29ed-4a17-a0b7-ec9dcf8c2b51 · outbound

This paper cites Near-optimal Regret Bounds for Stochastic Shortest Path.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Near-optimal Regret Bounds for Stochastic Shortest Path

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.101038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.101038Z digest=sha256:5559ed86fb1dc933969ad3eb6cc9a7bad17f378b61b3f229ed22201669c939b4

Observation 90004e48-3c74-4abc-a720-23e7249b81e6 · outbound

This paper cites No-Regret Exploration in Goal-Oriented Reinforcement Learning.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification No-Regret Exploration in Goal-Oriented Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.271979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.271979Z digest=sha256:81dcf21621a67ddab90a694f7e18eaf7cb499b35a5525fb5d994d2d9623c2ca0

Observation 33a6bffe-f7ac-42ee-bd1f-c8fb1e0d7025 · outbound

This paper cites an unresolved cited work.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.362134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.362134Z digest=sha256:620049bdd7529636492e46b5641be75a3ea747caff6d86d0c4b8487e69291186

Observation fc2fa487-25f3-47f5-b355-b9c98022b9fc · outbound

This paper cites 2007 , booktitle =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2007 , booktitle =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.540988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.540988Z digest=sha256:df54dd48a8a1fe219fbd34dd60747f1f632a4c2065ff71746c67159b06956e87

Observation 621bd805-52d6-48c3-9bc2-6bbb871c8922 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning, (ICML) , year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Proceedings of the 36th International Conference on Machine Learning, (ICML) , year =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.642174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.642174Z digest=sha256:d4ec4c553d7d632cadf01a56e3373eb484f8b0db8e15063391cdb17555507616

Observation c41b9d28-b731-4ba6-9a21-b7f8090d9d52 · outbound

This paper cites International Conference on Learning Representations , year=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification International Conference on Learning Representations , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.719137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.719137Z digest=sha256:be73956e82f914333c0c743b2aa8848dce1464fe84ea6ff1ce85565a18ac4f2d

Observation 54991ae1-8a39-4438-9cea-ee62dd14328e · outbound

This paper cites Kearns and Satinder P.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Kearns and Satinder P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.795125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.795125Z digest=sha256:8d42d1595589d64c76989d5b228fe80a5a406915277ed07ddc2a7b4b1f6fb7e0

Observation c1ce4b49-d576-4931-a150-6a86b5cf1ab5 · outbound

This paper cites 2012 , publisher=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2012 , publisher=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:51.871301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:51.871301Z digest=sha256:17689d642e34dcac6515370a263a24cbb2caa42b03d23ccf4e63efbf0adda55e

Observation 5ff6a3f0-f497-4930-a4f8-5b49a1dd56f7 · outbound

This paper cites Annals of probability , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Annals of probability , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.035309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.035309Z digest=sha256:a11ccede5c4e3d65348b2ae3b124fa18e2b302130c7963785ccc7a084cd562aa

Observation c017800b-d26a-44f1-849a-28bd475bf4b0 · outbound

This paper cites Proceedings of the 24th annual conference on learning theory , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Proceedings of the 24th annual conference on learning theory , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.197037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.197037Z digest=sha256:db23b80e90331712a9e1397e1d029ca9910bcd08fd6e22a77f1305a17b7c0e47

Observation 69df0cd3-7d10-48de-b2a2-6b3ea681f2c6 · outbound

This paper cites Annals of Statistics , Year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Annals of Statistics , Year =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.307664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.307664Z digest=sha256:8aee21c4ca531a030bd3dec6de9c30d316d08599a427f23573748fcf832298e0

Observation fce2aa21-e84e-4e75-972d-4bdc9cb0f8ab · outbound

This paper cites and Lattimore, T.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification and Lattimore, T

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.414397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.414397Z digest=sha256:8c45d85fdb6f23d47f8e054fd404e94e6df17efbf96c16a8beee8dfe81557a27

Observation 6be6006c-4c81-4d3a-aba0-c51c18fbbee8 · outbound

This paper cites Reward-Free Exploration for Reinforcement Learning.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Reward-Free Exploration for Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.528190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.528190Z digest=sha256:27372086c002ca190bb8010034673c4e2f8d2073efd6defc3ea79ce1d205862b

Observation 43b02cdd-8d43-41c1-83d9-e8b8bf164021 · outbound

This paper cites Jin and Z.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Jin and Z

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.631631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.631631Z digest=sha256:6e658c1ca36c34089195c704acbaa40313f17926b92141c2e3db46f2c8a6973c

Observation bcdc0c60-f423-4af0-9e75-bb72101a73fc · outbound

This paper cites KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.708972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.708972Z digest=sha256:30c29217293cc21b22fa60102221425ed13fd6de84458794d8984c96d40924c7

Observation 8d5d1644-bddd-439e-848d-db4ea399d3ab · outbound

This paper cites , Publisher =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification , Publisher =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.812900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.812900Z digest=sha256:88625c9423fbba9441ac8533aedef4acffc384e05656c3b564522002ef777dd7

Observation 619cd24b-501f-48f0-ad71-39b64b3191c7 · outbound

This paper cites and Li, L.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification and Li, L

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:52.928524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:52.928524Z digest=sha256:e6f1001301934735cdb1d92926e7da067f57af664f70254fa36fa5e26301c0a2

Observation 6075cc70-3582-42b6-8f2b-deb1491f643c · outbound

This paper cites , Author =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification , Author =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.039467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.039467Z digest=sha256:071667bf76c2c7d3be33e80675efa2100beabe5b49cb523b4c23899fb4aa35d9

Observation 3ce3eba5-58cf-4de6-a487-19278c927489 · outbound

This paper cites Planning in entropy-regularized Markov decision processes and games , booktitle =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Planning in entropy-regularized Markov decision processes and games , booktitle =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.144748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.144748Z digest=sha256:f4403e5b1326b54cc5eb9a534fce09f71bb16587e2e6d408bf85c58b1c125ec9

Observation b55058c5-aff3-4dad-a4ae-e970e9d1e40b · outbound

This paper cites Ajallooeian and Csaba Szepesv.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Ajallooeian and Csaba Szepesv

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.229216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.229216Z digest=sha256:e650d0147e27477ee86e2a91e80065ea8e4fa426adc572d6b09705386f015323

Observation c0a9dbdd-20d9-4a2b-86c7-a9d96be7a3f0 · outbound

This paper cites Koolen , title =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Koolen , title =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.390800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.390800Z digest=sha256:662e7345d8bef20b6a284109519419e31ea62e09dbe38326c7907ecded906330

Observation fb7479c0-ce41-41f2-8cfe-47d4277d7438 · outbound

This paper cites Proceedings of the 17th European Conference on Machine Learning (ECML) , Year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Proceedings of the 17th European Conference on Machine Learning (ECML) , Year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.545503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.545503Z digest=sha256:1581976ec75fdd1604c89dc1d8fd6861463b0f2387792d182ca2ab9e6a92419c

Observation 0ea14565-8891-409c-83c3-082e315e04ed · outbound

This paper cites and Mannor, S.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification and Mannor, S

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.659456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.659456Z digest=sha256:c89bd32e98331739d64e01e488310fd0c5b7ccdc386c3f9a4e28e2fe1981852e

Observation 00845ae3-c6ec-478e-8391-7362e1fe4a06 · outbound

This paper cites an unresolved cited work.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.768107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.768107Z digest=sha256:4a8c43888352517e1c8723721f43663f8c51b499104b16ccd30e3f49256f3725

Observation 979545b3-078a-48a1-8b45-08eea16e7008 · outbound

This paper cites and Cesa-Bianchi, N.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification and Cesa-Bianchi, N

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.873483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.873483Z digest=sha256:339e430656f87871a444c435f73a2633353598f9c93aabbeca165b78c50752c3

Observation 4e6fb36e-d87f-449e-add8-de587bbbc96b · outbound

This paper cites Jaksch and R.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Jaksch and R

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:53.937202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:53.937202Z digest=sha256:cb3eca52d380ae170396ffa97792fb5da62e42ca25ecb56023243e5eb6cd8190

Observation 189483f2-9f9a-4678-ae46-335939591b2b · outbound

This paper cites Proceedings of the Twenty-Sixth.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Proceedings of the Twenty-Sixth

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.026928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.026928Z digest=sha256:9f17967a048812bda990b42a7cb494c89fdf26f3f7c763c91876871b863484b8

Observation fff355dc-4095-4272-8b56-3184e50eede6 · outbound

This paper cites an unresolved cited work.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.126001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.126001Z digest=sha256:fafaa122343db3f151827cb7bc48833322973138e0d8ffd844e55607adfd25bb

Observation 907537d2-b61e-4f19-9ad6-bac4d07e8669 · outbound

This paper cites Journal of Artifial Intelligence Research , volume =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Journal of Artifial Intelligence Research , volume =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.180380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.180380Z digest=sha256:e2e518ae752ba8d226a17f11b48bdf75f1585db583c8101c38c0dec9e987a1d3

Observation 91f19420-af17-4b5a-935f-0d86275e7478 · outbound

This paper cites Kearns and Yishay Mansour and Andrew Y.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Kearns and Yishay Mansour and Andrew Y

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.262136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.262136Z digest=sha256:f43c69077c416ac2080b02e292983d68277335f29bee734f3f6adf9eb3e6aca0

Observation d125c550-1598-41fb-8fdc-db510f88095e · outbound

This paper cites Advances in Neural Information Processing Systems (NIPS) , Year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Advances in Neural Information Processing Systems (NIPS) , Year =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.330410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.330410Z digest=sha256:05225a4e0f30a2075619192b2efb84e5c7b2019cdb6986f8b2e79407f6cf72c1

Observation afe67170-118d-4b0c-98de-b4c10abe8635 · outbound

This paper cites 1998 , Owner =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 1998 , Owner =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.424017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.424017Z digest=sha256:c1d9a464f22c68d3d037f8e197dd7b57a97e5af8a2eb9ae8df7059d518076f7a

Observation 9e05cb1e-1dc3-4b9d-b17a-c1aae01bbdb6 · outbound

This paper cites Lillicrap and Karen Simonyan and Demis Hassabis , title =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Lillicrap and Karen Simonyan and Demis Hassabis , title =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.524174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.524174Z digest=sha256:b1ffbd207084dc15048e3e582e43122bb68d0a462b97a4c6b213c0b408cb462b

Observation 72641798-1f03-4c03-86ce-b3c0ebabff9a · outbound

This paper cites Advances in Neural Information Processing Systems (NIPS) , year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Advances in Neural Information Processing Systems (NIPS) , year =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.624236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.624236Z digest=sha256:8a291e678b53881dcd34c0dc20ede87d51b6909caa6e811fe1529670a0ed570b

Observation 62278889-d502-4f51-bcce-15aa8724657f · outbound

This paper cites Jonsson and E.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Jonsson and E

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.716414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.716414Z digest=sha256:15214d13a39bd64c1d114cd6485a30047e9229c3a865b62c1b1d46ecbee71a7b

Observation b782cc45-4b09-48b5-b09a-6a3f84248914 · outbound

This paper cites Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.826178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.826178Z digest=sha256:56c77d5e52b537b7ca640e5c473bfe5adda10282652f6311dea492f1e05effd2

Observation 2bc6c379-9148-4c9a-aa56-2e07992e5c03 · outbound

This paper cites IEEE Transactions on Computational Intelligence and AI in games, , Year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification IEEE Transactions on Computational Intelligence and AI in games, , Year =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:54.975080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:54.975080Z digest=sha256:9a0b1477167e86e9e1df5dbbf1e8b47ff12d8be44a7fbe5954e11dbfbcd1bc6d

Observation 9ae019b0-8e74-442f-a871-426d26ce8dea · outbound

This paper cites Neural Information Processing Systems (NIPS) , Year =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Neural Information Processing Systems (NIPS) , Year =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.117148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.117148Z digest=sha256:402744d1c794f82a8021c43baa45e8757e743e05d418f5992c4ac39690eaad00

Observation a55a6404-fccd-47f2-a36c-ac15caa3d474 · outbound

This paper cites Advances in Neural Information Processing Systems , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Advances in Neural Information Processing Systems , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.204020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.204020Z digest=sha256:59f94c2072a28deb184ce9378ed2bf34e407aef81a393af5db5587da25160611

Observation a85bf99e-c40e-40eb-b99a-81518bea411f · outbound

This paper cites Gheshlaghi Azar and I.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Gheshlaghi Azar and I

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.321130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.321130Z digest=sha256:b39c765860aeb59aaa46f435bb1fbf75df1ce62446ed718aacc4d44e856a11bd

Observation 625e60b7-0abe-408e-8596-367787ecafc3 · outbound

This paper cites On the Sample Complexity of Reinforcement Learning with a Generative Model , booktitle =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification On the Sample Complexity of Reinforcement Learning with a Generative Model , booktitle =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.502056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.502056Z digest=sha256:80c2e13efd4df32676d5f879f417b5f4a4d0f29e70220c4103a85ab869cfaed2

Observation 2d59a362-9d72-43da-be6a-d6e0001eb11e · outbound

This paper cites Kearns and Satinder P.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Kearns and Satinder P

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.676277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.676277Z digest=sha256:e30057216c40a4c9960ae54d6bc8a19986cb22cb8f901e1baed8db685442401d

Observation 74961bd1-8f15-4d53-aff5-ed0f843ba1ba · outbound

This paper cites Efficient Reinforcement Learning , booktitle =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Efficient Reinforcement Learning , booktitle =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.819981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.819981Z digest=sha256:dc7f2457234323d9efda4d008abfdc1de1856e862b04237a180b1f6403c0095c

Observation 53b09fb9-4a99-4d98-adc9-17aea1cdb6bc · outbound

This paper cites Expected Mistake Bound Model for On-Line Reinforcement Learning , booktitle =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Expected Mistake Bound Model for On-Line Reinforcement Learning , booktitle =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:55.933964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:55.933964Z digest=sha256:f6f2324f18ab11a960a18ac43cb5edd43e61af6327e869936bc849f889f0f1df

Observation 1fbd94d1-4ec0-4d2e-9704-c101a36f98b7 · outbound

This paper cites and Capp.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification and Capp

Reference 52

Resolution
parse uncertain
no resolver link, observed 2026-08-03T09:55:56.111870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.111870Z digest=sha256:a972559d9297ba171dffc29f636743634e1945b2346d779b09f087eea2bed291

Observation 5fa62a19-2191-4e92-b73e-155fa67f489a · outbound

This paper cites Brafman and Moshe Tennenholtz , title =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Brafman and Moshe Tennenholtz , title =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.223829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.223829Z digest=sha256:bdb776aea16d7cc77adf7a0e6664b0bbc8c5d2d81c7256e72cdd8a32d6d0c12b

Observation a94b3860-342b-47d1-9674-74416643cec2 · outbound

This paper cites Strehl and Lihong Li and Eric Wiewiora and John Langford and Michael L.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Strehl and Lihong Li and Eric Wiewiora and John Langford and Michael L

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.377054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.377054Z digest=sha256:b8277b3e5d43a72b20c4e4180e536a11fadf50a8ab5a7cf791917eaa55fa752a

Observation 67f7d795-f4f5-4231-a502-cf2e465db8a0 · outbound

This paper cites Strehl and Michael L.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Strehl and Michael L

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.520976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.520976Z digest=sha256:74d3793c2142e0a665ab6550f67c33c24d2094e0d47e62d7f629f0cface0f870

Observation e28f6bc2-3415-480f-bc9b-ace9ac5eeb01 · outbound

This paper cites Conference on Learning Theory , year=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Conference on Learning Theory , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.570788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.570788Z digest=sha256:c7f2a6f27012c6786a83d8d25ccbccabdf7222b14192e8f25a63601768da1531

Observation 96451d14-88c1-4a6e-a7e3-438df515c703 · outbound

This paper cites 2019 , booktitle=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2019 , booktitle=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.654008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.654008Z digest=sha256:50bff08e0d18a674bcdde069b702483bb16e35b6fd6ea5b21fc0317e1d1de0d0

Observation 022977d2-9f63-474d-ae51-e5ce28d6baa0 · outbound

This paper cites Artificial Intelligence and Statistics , pages=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Artificial Intelligence and Statistics , pages=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.711295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.711295Z digest=sha256:87fac16f7c3e600a642d6ca45ce250398dc0d93867aec4f33f7ab0ed06f14911

Observation a80c9e25-cba2-44d4-8540-be95e69f3f27 · outbound

This paper cites Advances in Neural Information Processing Systems 28 , editor =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Advances in Neural Information Processing Systems 28 , editor =

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.791193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.791193Z digest=sha256:eaed4d4a6a24d4d2637a054f13ce829318c1861fe5ae6be2be47f55e4861c8ac

Observation 726e78c2-0162-44a0-ac72-6ffb0a0af011 · outbound

This paper cites 2016 , eprint=.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2016 , eprint=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.847218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.847218Z digest=sha256:199e33d78b9870888750c3ac616096403febec4a89aa526eaabbe0f4cf4040a2

Observation 4672fa2e-fc9b-4e3b-ac39-90f2dd644e10 · outbound

This paper cites 2012 , Journal =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification 2012 , Journal =

Reference 61

Resolution
verified exact
doi, observed 2026-08-03T09:59:01.344957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-03T09:55:56.913323Z digest=sha256:6bc6c9ccc14948561538b6159481c1565ed383ca2ef8cd9c835deab4ebf8b1b3

Observation 91c50cc8-3242-4622-9b1a-cb5ae2daf2d4 · outbound

This paper cites Advances in Neural Information Processing Systems 17 , editor =.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Advances in Neural Information Processing Systems 17 , editor =

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:56.939347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:56.939347Z digest=sha256:95019a87aa3ce4e5dbbdf9e3762b6a4db2bf4edc78323fdfbe1c9cb59b163915

Observation ee2d9232-b4ba-489f-8d9f-c54ebef1d4d5 · outbound

This paper cites Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.097421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.097421Z digest=sha256:eac79e503328cb3268a6101b8a8a3f4aa86e9cfe3e1115fab718db00e0f88f4e

Observation 8dd8c9bd-b5e0-4a8d-b5e2-39da940a405c · outbound

This paper cites Kaufmann and P.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Kaufmann and P

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.157413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.157413Z digest=sha256:357535c653d47838615454a1c62891c1292908b199c7021777034420d148c0b4

Observation c9bc94a5-5d60-468c-9fbe-9192cee7815b · outbound

This paper cites M\'enard and O.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification M\'enard and O

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.301408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.301408Z digest=sha256:4df3968c6c4f949753f76f5e45872e97cce6c672e2a8d75878a25ce75544cf4e

Observation 2cf5ecf5-ab5b-4749-bfd3-d9ec8a5426e2 · outbound

This paper cites Infante and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Infante and A

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.469624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.469624Z digest=sha256:f2d3d9f770200a3606379217dd635bb55daff547ddbf60861bdd61b9f5183193

Observation d2db58e5-94fd-4d64-87dd-8612ad633fee · outbound

This paper cites Drappo and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Drappo and A

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.635393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.635393Z digest=sha256:ab1a260b15d250c1a5c92a6134d0300a38131711619c5f0ec5e82a62bbfc3e0b

Observation 48fb65db-f93d-41a0-b528-654501251544 · outbound

This paper cites Robert and C.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Robert and C

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.732619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.732619Z digest=sha256:8c9048234a2e2ed0d2afabdd638bd71f784fc7d43e011e587bee2984189a01c6

Observation 1e6e4e3b-760e-484f-8e38-ed5e1192d063 · outbound

This paper cites Wen and D.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Wen and D

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.844780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.844780Z digest=sha256:2ae3cd1546fd71b77161086f0b2e90c4dd8b779c724b6573e12fffa3c349be36

Observation 3f445232-6a50-4553-9f44-cef3a98d419a · outbound

This paper cites Nachum and S.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Nachum and S

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:57.953298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:57.953298Z digest=sha256:8509af9a2468db51535a5667e1329b0c7392897e49ad734f57f61dc24ee7a9c8

Observation ad5006ea-0472-4f4d-98a5-93874f136765 · outbound

This paper cites Konidaris and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Konidaris and A

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.032888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.032888Z digest=sha256:d1146c84c6bacd1524802139b6eaffeaad2a24db28cdaa566e8cafb81e7fa51c

Observation 837d335b-6f32-454d-af8c-847122eef906 · outbound

This paper cites an unresolved cited work.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.107434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.107434Z digest=sha256:58a3a00a7d8bcecb69403c768a9706f2023fd19302e58524da6d5e5e77ca9ecb

Observation b89c28c5-94b6-4b18-93b3-bc432e9cf0cd · outbound

This paper cites Levy and G.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Levy and G

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.240938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.240938Z digest=sha256:f80aee348334ee7964b62fce5e04bc144ff2aaa5f895dbfa444b334e5f842d9d

Observation 06e71de6-a0ca-4a0c-a92d-531379a69f41 · outbound

This paper cites Fruit and M.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Fruit and M

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.318375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.318375Z digest=sha256:d510385afa2964b5e335c9b08d0148b287bd0a137fed010b8216cc806d4a988a

Observation 1d2448fa-d25d-41b8-8dc4-40e86259aaa9 · outbound

This paper cites Drappo and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Drappo and A

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.448196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.448196Z digest=sha256:f02e6d0864942fc725085a518b8a336f3d6c894ff8903a6fa587861aa68c40cf

Observation 7a04f65b-c566-46ad-9d29-5d3c8af77afb · outbound

This paper cites Rafati and D.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Rafati and D

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.620135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.620135Z digest=sha256:0114c09261c94b9f231818b5cc65d095d5433d62953bce8f059a299617c7973f

Observation 53097f2b-3b9c-4d69-9653-42e1bf2300ad · outbound

This paper cites Gopalan and M.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Gopalan and M

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.738145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.738145Z digest=sha256:c4a170d9bc010139be27b10fcb7af1e0824a6fc83c6f27ddb93d600a0bd199cf

Observation aad61cbc-29eb-4997-a6da-a5f4009514d1 · outbound

This paper cites an unresolved cited work.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.807958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.807958Z digest=sha256:0bdb4f0efe47241f600dee2ac33b6adddb873d2752d4eecc6b1356cf03298d1d

Observation 54098ecd-5095-4b60-8329-20be4a5b9eab · outbound

This paper cites Ahn and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Ahn and A

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.877754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.877754Z digest=sha256:87e097a5d3c1fd6668bc92e718d9f8683ee8cd1d91c5dfe4cbfbd7e73099f9df

Observation c1904535-66ed-4613-b335-daec52f79b0d · outbound

This paper cites Brunskill and L.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Brunskill and L

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:58.943364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:58.943364Z digest=sha256:06ca184ef4e48022fc04524c0763f6f31f3ef3f70d52ca718f3f98d401d96384

Observation 97155275-ee0e-4665-9425-6cf60a68d57c · outbound

This paper cites an unresolved cited work.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:59.105372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:59.105372Z digest=sha256:399139bab65a990c021ef869aaaa3baad5d19e26a86e330028e2235d2477880d

Observation 63857bf5-7e2a-4e87-80e2-a06c35ee0201 · outbound

This paper cites Matthews and M.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Matthews and M

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:59.235825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:59.235825Z digest=sha256:fd81df627de7aa8fe5a99a0fa8cc3e9c03ada55ec1d79ce038810a47424cf249

Observation f2e6f85f-512e-4175-8630-495d3f910034 · outbound

This paper cites Drappo and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Drappo and A

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:59.386774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:59.386774Z digest=sha256:1f4878bbf8753c5c165f07946338cbdbb350d882163a36b36f43fb4488f365a6

Observation 90f241a2-765e-4fef-be3b-393448182db3 · outbound

This paper cites Kuric and G.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Kuric and G

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:59.534752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:59.534752Z digest=sha256:11d19847838d0f34db75122525157cc83ac7c0b3b1daf3071a0158a6d51db35f

Observation 52f7b743-e91e-4618-8713-36cbb4e8115b · outbound

This paper cites Manenti and A.

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Manenti and A

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T09:55:59.660672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:55:59.660672Z digest=sha256:36e690e2f4d6fd06a52c533388eef0db1ceb2fd9c506b760d2ec0482df0581d5

Pith citing papers

No inbound Pith citation observations are available.