Pith. sign in

Paper Citation Record · LEDGER

Provable Partially Observable Reinforcement Learning with Privileged Information

As of 15 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2412.00985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00985 v3

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:00:57.747243Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact1
  • verified fuzzy67
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7163470-73cd-472d-b4a6-f1ebc205c09e · outbound

This paper cites End-to-end training of deep visuomotor policies.

Provable Partially Observable Reinforcement Learning with Privileged Information End-to-end training of deep visuomotor policies

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.378425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.378425Z digest=sha256:c379f8c58bf79d80fe4e1dd6eb964dd87dd0f8429628aac7f93f5d18b4ab0921

Observation 5cbeeea6-e844-492d-99e6-5820f41ad289 · outbound

This paper cites Learning dex- terous in-hand manipulation.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning dex- terous in-hand manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.383149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.383149Z digest=sha256:3db3b045b1994d05b2179948295b9f9304a3295568d4890d0edf9382189bf2b8

Observation 1dd38f47-686d-40a7-a6b4-632b05f0de02 · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Provable Partially Observable Reinforcement Learning with Privileged Information Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.387345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.387345Z digest=sha256:5634999e53d9b60de957d5a55b9031d5baacfe8eeabc416aad2139a4c1d5922e

Observation ea9f9c37-b818-4012-99d4-d365c7d89a29 · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey.

Provable Partially Observable Reinforcement Learning with Privileged Information Deep reinforcement learning for autonomous driving: A survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.392093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.392093Z digest=sha256:153a391f80981d420cb0611779501b9eb332dce24730e45c0d5085446211ce8f

Observation 6cd51fa3-3e5b-4ea3-b2a2-72417c621d80 · outbound

This paper cites Pomdp-based statistical spoken dialog systems: A review.

Provable Partially Observable Reinforcement Learning with Privileged Information Pomdp-based statistical spoken dialog systems: A review

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.396226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.396226Z digest=sha256:2173e57f8c47698c36916edb0caed4fce54af95caffb404f18f43bc60c8d867f

Observation c90bb442-a7bd-4b7e-8cb8-45907d5e3881 · outbound

This paper cites Informing sequential clinical decision-making through reinforcement learning: an empirical study.

Provable Partially Observable Reinforcement Learning with Privileged Information Informing sequential clinical decision-making through reinforcement learning: an empirical study

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.400361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.400361Z digest=sha256:7a6da6ebabebd2cef621387d3ff97c969627fb912ee98b99147cf39995b1005c

Observation a70d1743-7a2a-4167-b3e5-3c6ae7eb9f07 · outbound

This paper cites The complexity of markov decision processes.

Provable Partially Observable Reinforcement Learning with Privileged Information The complexity of markov decision processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.404638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.404638Z digest=sha256:a449c7f96f6d7cad2a608e6a4fa4da72e988aaf7cc60d0d8e36908a0fddc0623

Observation 23b1aab0-39eb-4063-bc2a-0f7b5c539fe7 · outbound

This paper cites Pac reinforcement learning with rich observations.

Provable Partially Observable Reinforcement Learning with Privileged Information Pac reinforcement learning with rich observations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.408509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.408509Z digest=sha256:ade648beead2c0531f18b8f5f156f87cf32f4cba24f6920123c76196ace0019f

Observation 6edf07d0-4530-4dc5-9c96-d945087cfccc · outbound

This paper cites Sample-e fficient reinforce- ment learning of undercomplete POMDPs.

Provable Partially Observable Reinforcement Learning with Privileged Information Sample-e fficient reinforce- ment learning of undercomplete POMDPs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.412288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.412288Z digest=sha256:e315f400e62e6241a76c87f1b4a2fdd68df82b0067aa1d6fd32a8df5af5db79a

Observation 35341c44-0538-44a0-ad0d-25693ee815aa · outbound

This paper cites A counterexample in stochastic optimum control.

Provable Partially Observable Reinforcement Learning with Privileged Information A counterexample in stochastic optimum control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.416105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.416105Z digest=sha256:d49c6e2cf531b24fbf3a37eb6e6c56355aefca3c2309acd7c6f8e569c9820a2f

Observation 85cdf6c5-84e0-402f-add6-a0d965b262bd · outbound

This paper cites On the complexity of decentralized decision making and detection problems.

Provable Partially Observable Reinforcement Learning with Privileged Information On the complexity of decentralized decision making and detection problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.419988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.419988Z digest=sha256:560930c005a0249e5d315f9dcdce728ff26ed5e37bcbfa3d3814056113f20a0c

Observation 0e9392e0-05ed-48ab-a20b-8989b49afb6f · outbound

This paper cites Multi- agent actor-critic for mixed cooperative-competitive environments.Advances in Neural Informa- tion Processing Systems, 30, 2017.

Provable Partially Observable Reinforcement Learning with Privileged Information Multi- agent actor-critic for mixed cooperative-competitive environments.Advances in Neural Informa- tion Processing Systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.424233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.424233Z digest=sha256:15377339df165e97f26078d7947b8c2d1a88ea2b9a3eb26fb8020e534cea3426

Observation 4f837fb5-a935-4d6e-b68f-d7ec928e50b6 · outbound

This paper cites QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.428099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.428099Z digest=sha256:25def852254cd0f460c3e655088c0b9895a5e09b263326a8221c8eaf4b7b1368

Observation 0b6db165-42da-4fcc-b2dc-b38653b2e6c9 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Provable Partially Observable Reinforcement Learning with Privileged Information Counterfactual multi-agent policy gradients

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.431857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.431857Z digest=sha256:a18d9f138b34a1acb09be0290c520353f39f990fac7b7bcdfcc40a9350d8921c

Observation 4ae1370d-cfe2-4e80-9d80-3e49b8411cd9 · outbound

This paper cites Grandmas- ter level in StarCraft II using multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Grandmas- ter level in StarCraft II using multi-agent reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.435449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.435449Z digest=sha256:e0bd7bec08ab01cd9829f7b7d5f1345a90fcfe4a72a8e204bdccbc84d4b19839

Observation 209f5775-dff3-4ee1-8fbc-63ef7fd0f64c · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning quadrupedal locomotion over challenging terrain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.439379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.439379Z digest=sha256:675c3e3ad57706f0ababe98b9206710c142eb3c6e4b883b75f8757bbc2c5c23f

Observation b086a5a3-43c3-4730-9c9c-83b87f90fc3f · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.877864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.443169Z digest=sha256:7fa0150716f7593b392704cc2d0a9ce6d2037286eac80db5aa8e2f051745dc7d

Observation 72e2ad17-e185-48ce-962d-7f900b515742 · outbound

This paper cites Learning by cheating.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning by cheating

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.865509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.446983Z digest=sha256:26af8490f302aa5b24f6a2a07bab9599b9a124073373e107fd6eeffe1eb40e31

Observation f8827aa1-a12f-4b16-b09e-b5af8efc6b35 · outbound

This paper cites Asymmetric actor critic for image-based robot learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Asymmetric actor critic for image-based robot learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.853512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.450732Z digest=sha256:88071fdfcf27015cf3a158b978d87a062c3e255d429a0928e31b93dd8c3289bc

Observation 3dbd4a4d-4b10-4558-9917-3494ab24f717 · outbound

This paper cites Learning in pomdps is sample- efficient with hindsight observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning in pomdps is sample- efficient with hindsight observability

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.841243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.454380Z digest=sha256:f049ad01f8ece82b051a585c403969a8d8f4b3bff77c38cd7225579726244355

Observation 4833688b-eb7c-41d0-a7a8-da1d7b233706 · outbound

This paper cites Sample- efficient learning of pomdps with multiple observations in hindsight.

Provable Partially Observable Reinforcement Learning with Privileged Information Sample- efficient learning of pomdps with multiple observations in hindsight

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.828703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.458190Z digest=sha256:9700abb07e6f43b04a10bc45135757ca0143c561022582c1cb0a682921b8a382

Observation afb14ff5-9add-461f-830f-dd4432bae525 · outbound

This paper cites Learning in observable POMDPs, without computationally intractable oracles.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning in observable POMDPs, without computationally intractable oracles

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.815988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.461960Z digest=sha256:ccf1966766c831c558b6e61487e27e90580a4e408b41a7d8bf171102b1cd3af2

Observation 6b6818fe-b22c-491d-854c-f95eb52a8bed · outbound

This paper cites On oracle-e fficient pac rl with rich observations.

Provable Partially Observable Reinforcement Learning with Privileged Information On oracle-e fficient pac rl with rich observations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.803818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.465780Z digest=sha256:12f741905802a69f3bbc84012e59ba015f045b883e665926d76a80ea1459f21a

Observation 47c406f7-1a2d-4fae-afdd-5786d14ebff9 · outbound

This paper cites Provably e fficient rl with rich observations via latent state decoding.

Provable Partially Observable Reinforcement Learning with Privileged Information Provably e fficient rl with rich observations via latent state decoding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.791388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.469539Z digest=sha256:13721b45114a69e6dc1969c482a7abebe365682f405408661a0dae99eaae0f35

Observation cc6b3e90-541e-4674-8d86-9ce17c4e78f7 · outbound

This paper cites Kinematic state abstraction and provably efficient rich-observation reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Kinematic state abstraction and provably efficient rich-observation reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.779272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.474307Z digest=sha256:f4ec3d014dc2d901e0f8d8d4b25664b5cee3ecadb7d6f4b694dd618c91dc590d

Observation 217b5634-c3f1-4990-989d-548706502368 · outbound

This paper cites Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:00:57.892712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.478097Z digest=sha256:db95e4a51c60a37cc4f90c462e59442a238148121e43152b49a7aa372ce6d325

Observation d6220eb8-cd38-4b8b-b43c-fb142e56c383 · outbound

This paper cites Provable reinforce- ment learning with a short-term memory.

Provable Partially Observable Reinforcement Learning with Privileged Information Provable reinforce- ment learning with a short-term memory

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.766737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.482113Z digest=sha256:f78db36b600562c4aaa23507d0306c4eb2720da479c313bbc4424788a38dc12f

Observation cd64147e-b6ea-4d4e-9728-8d54b405e90e · outbound

This paper cites When is partially observable rein- forcement learning not scary? In Conference on Learning Theory, pages 5175–5220, 2022.

Provable Partially Observable Reinforcement Learning with Privileged Information When is partially observable rein- forcement learning not scary? In Conference on Learning Theory, pages 5175–5220, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.754554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.486135Z digest=sha256:33504353e8f5e424ccb015b3f333398dc0568cad6986cf100fc4ba29ec9092d7

Observation 737e6c1e-7190-412b-909f-d6d980fe8e94 · outbound

This paper cites Partially observable multi-agent RL with (quasi-)e fficiency: the blessing of information sharing.

Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable multi-agent RL with (quasi-)e fficiency: the blessing of information sharing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.742068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.489987Z digest=sha256:c2200499a6dd6a0a7e1910272e7083a2a3d02780383e6fb272b4f4a91ca71d1c

Observation c383ef95-b6ce-471b-b251-94451ae6aab7 · outbound

This paper cites Schapire.

Provable Partially Observable Reinforcement Learning with Privileged Information Schapire

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.729516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.493807Z digest=sha256:f5e341fd4adba6012b4b2dc7cd5b9ede0cab16e3951c0234453b4f55ebecd931

Observation 48aa0392-b631-4fd2-89ff-f49f5db7b751 · outbound

This paper cites Represent to control partially ob- served systems: Representation learning with provable sample efficiency.

Provable Partially Observable Reinforcement Learning with Privileged Information Represent to control partially ob- served systems: Representation learning with provable sample efficiency

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.717368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.497654Z digest=sha256:dfeafdb648a5e2c681f1de14173f0fa77a167c549fa708b0fb17811a2fb49447

Observation e2754bc8-d010-4a96-be31-39ae8a9d93a3 · outbound

This paper cites Partially observable RL with b-stability: Unified structural condition and sharp sample-e fficient algorithms.

Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable RL with b-stability: Unified structural condition and sharp sample-e fficient algorithms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.705076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.501395Z digest=sha256:76b5cc92e1399686bd217b0a4add17e6c3123d028c4e5aa0c40c54d0e97f895e

Observation af9dad91-3fc7-48d2-89f3-f5753195c22f · outbound

This paper cites Reinforcement learning from partial observation: Linear function approximation with provable sample e fficiency.

Provable Partially Observable Reinforcement Learning with Privileged Information Reinforcement learning from partial observation: Linear function approximation with provable sample e fficiency

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.692309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.505120Z digest=sha256:f034c501a67fcffa68e2e512739baa8b17cb2f106be41e1c9be5dbc8aaf67764

Observation b66ce82c-c450-40e6-93c4-2050fb6e6a4f · outbound

This paper cites Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision pro- cesses.

Provable Partially Observable Reinforcement Learning with Privileged Information Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision pro- cesses

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.679947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.509825Z digest=sha256:b4d364e63a08300a959dab9dd1a2619d462e33035d7fb6d437a9141fb7641cd9

Observation aec8ff2d-a16a-497f-9e24-2e5ac3d3718c · outbound

This paper cites Optimistic MLE: A generic model-based algorithm for partially observable sequential decision making.

Provable Partially Observable Reinforcement Learning with Privileged Information Optimistic MLE: A generic model-based algorithm for partially observable sequential decision making

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.667308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.513652Z digest=sha256:5c213edacc7d3342d0b756f541e54b692bfae00a199f17bf2fac02df6e7da1c7

Observation 86df68e4-ed4d-4a4f-8aa1-6c371f8fe8fd · outbound

This paper cites an unresolved cited work.

Provable Partially Observable Reinforcement Learning with Privileged Information Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:00:58.655043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.517747Z digest=sha256:748eaa3bf75aa645ab9c9796f2ae5ef7ea3ad41a96e06afe82248c67faa25045

Observation 72599437-b993-4eb5-a50e-1f38ec65ac56 · outbound

This paper cites Partially observable multi-agent reinforcement learning with information sharing, 2024.

Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable multi-agent reinforcement learning with information sharing, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.642424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.521399Z digest=sha256:ad99e947bb13f95182d122256144fc972978e30ffb707515c69fe2b358ce2edb

Observation 105fcce5-289a-46c5-bcf2-969200142602 · outbound

This paper cites Planning in Observable POMDPs in Quasipolynomial Time.

Provable Partially Observable Reinforcement Learning with Privileged Information Planning in Observable POMDPs in Quasipolynomial Time

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.525271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.525271Z digest=sha256:3bc49d1082c9a941c5f22d47e2a3ac5a820ca038828049d9d65c72a989a124d8

Observation 8cb4df1d-8d25-4114-9ec0-02605981450a · outbound

This paper cites Planning and learning in partially observ- able systems via filter stability.

Provable Partially Observable Reinforcement Learning with Privileged Information Planning and learning in partially observ- able systems via filter stability

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.629812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.529316Z digest=sha256:b1a0d41fd676dd32c0da2405f12c03d8f77ef47f69ba0672b2b7be452d28e0f3

Observation 4b795815-13fc-490f-a582-00a8188a81bd · outbound

This paper cites Theoretical hardness and tractability of pomdps in rl with partial online state information, 2024.

Provable Partially Observable Reinforcement Learning with Privileged Information Theoretical hardness and tractability of pomdps in rl with partial online state information, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.617712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.533144Z digest=sha256:6e7c52e44af00857b33c19f8a599f70e12d60f928d620cd269b3c14cfb46a0c7

Observation 311d25fe-b0d2-4841-a9de-df18d7812141 · outbound

This paper cites Leveraging fully observable policies for learning under partial observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Leveraging fully observable policies for learning under partial observability

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.605312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.536883Z digest=sha256:7c08f8136f0a4b9982e8ad583d56a3898a8198e3a8b4a55735f54b10b8c803a9

Observation 8343e49c-fd3a-448d-994e-3cd4802dda99 · outbound

This paper cites Learning to jump from pixels.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning to jump from pixels

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.593020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.540594Z digest=sha256:c7e802318e8b569b0029554f8b142c5750559e7b1a870980eaa19d07dc4aeb9c

Observation dca28fe8-bbb3-4f3e-9302-f672302aeb71 · outbound

This paper cites Tgrl: An algorithm for teacher guided reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Tgrl: An algorithm for teacher guided reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.580773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.544380Z digest=sha256:6f20aef732f1ba3f5e9b1d7b3ef77deecfb217ecbebff38f0fcbd1ce00e18942

Observation 27473edb-659a-4496-8580-a62aae1ad57c · outbound

This paper cites Asymmetric DQN for partially observable reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Asymmetric DQN for partially observable reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.568365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.548314Z digest=sha256:3dff98171e7d4c02cd7f9d2da6e01134b55c0aad3b8408d268762a6630bff082

Observation 73818330-0008-423b-82cd-5428279c758f · outbound

This paper cites Perfectdou: Dominating doudizhu with perfect information distillation.Advances in Neural Information Processing Systems, 35:34954–34965, 2022.

Provable Partially Observable Reinforcement Learning with Privileged Information Perfectdou: Dominating doudizhu with perfect information distillation.Advances in Neural Information Processing Systems, 35:34954–34965, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.556140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.552052Z digest=sha256:5ebc820f92dc578cd1b0831b3e5a87c19b23fa91aeb1c8a7654186b511084cb8

Observation 86c05e3a-ba6c-4130-ad4c-fc2068feb608 · outbound

This paper cites Towards unifying behavioral and response diversity for open-ended learn- ing in zero-sum games.

Provable Partially Observable Reinforcement Learning with Privileged Information Towards unifying behavioral and response diversity for open-ended learn- ing in zero-sum games

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.544002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.555842Z digest=sha256:6348447519bb36a6cf0d2486a699c98936b0b5a0beece227e19abcc32ec72458

Observation 46cfd868-9448-4408-8a6d-965ecd08c15e · outbound

This paper cites Unbiased asymmetric reinforcement learning under partial observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Unbiased asymmetric reinforcement learning under partial observability

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.531166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.559621Z digest=sha256:5495121fad785d2002456c5cc594ff8bcf092c7c321fb6b7bca8acf26737bf48

Observation d18c399a-2656-4fb1-9490-1eb289e362f7 · outbound

This paper cites A deeper understanding of state-based critics in multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information A deeper understanding of state-based critics in multi-agent reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.518806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.563346Z digest=sha256:186b3f6f71d484229a25ee5b2ff934ce73da48fd4409ef9f226dcc634ea928e3

Observation 992bd910-ae36-4806-9b81-d38c65da0108 · outbound

This paper cites On cen- tralized critics in multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information On cen- tralized critics in multi-agent reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.506483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.566988Z digest=sha256:28f43cb4af2c65e14a8e764c3c51af9b353a75ae7eb84267100792db23c23cc0

Observation 2e1a4770-5d02-4db9-b30e-1e7d478fe43c · outbound

This paper cites Learning belief representations for partially observable deep rl.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning belief representations for partially observable deep rl

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.493910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.570730Z digest=sha256:7f22c462b0cedd249ac648fd9736a05cd2b5277a120d6df83a2161c46919aecd

Observation c202d67a-d224-46c0-8f92-717a8020d41d · outbound

This paper cites Learning belief representations for imitation learning in pomdps.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning belief representations for imitation learning in pomdps

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.481653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.575636Z digest=sha256:25c2ce545bbb1db25167ed7ce6d096b50a7bffc91b71e765b10c2fd383764faf

Observation bf86327d-f6bd-452b-95fd-7fd78af64a55 · outbound

This paper cites Belief-grounded networks for accelerated robot learning under partial observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Belief-grounded networks for accelerated robot learning under partial observability

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.469046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.579698Z digest=sha256:c19c31cbe04d869f00997ba8eb2eeb7c911c0d56da11b860ab02f88690eab357

Observation 1b8675e2-1378-493b-9b45-0aae522d1133 · outbound

This paper cites Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings.

Provable Partially Observable Reinforcement Learning with Privileged Information Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.583507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.583507Z digest=sha256:5f293673b86842dc53e7956fec01000f50ac9ff27da24ad2d266073dfa6b4891

Observation 892cabc8-9d4a-43ba-b625-39329c5c8d51 · outbound

This paper cites Flow-based recurrent belief state learning for pomdps.

Provable Partially Observable Reinforcement Learning with Privileged Information Flow-based recurrent belief state learning for pomdps

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.456809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.587635Z digest=sha256:11218c0ec4b42614569cc27050128df981bb5d560562a74ed51610a12b955d1d

Observation 4084d94e-4e65-4da7-8968-b4cec9add4ab · outbound

This paper cites Belief state actor-critic algorithm from separation principle for POMDP.

Provable Partially Observable Reinforcement Learning with Privileged Information Belief state actor-critic algorithm from separation principle for POMDP

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.444516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.591604Z digest=sha256:83db0dd2a89352aaadef6b1a2f421a42e9fd8e30e59a3d1b7637b70467280a8a

Observation 2882b484-5216-4df3-82a5-5405accbbe24 · outbound

This paper cites Neural belief states for partially observed domains.

Provable Partially Observable Reinforcement Learning with Privileged Information Neural belief states for partially observed domains

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.432094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.595393Z digest=sha256:2466e4e002bf0dc7abdd97bed49e85181849dd15b66aab287c0841eeff47d414

Observation 6691a7d0-4711-4ad8-b0b7-68895a1f9f99 · outbound

This paper cites The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models.

Provable Partially Observable Reinforcement Learning with Privileged Information The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.419475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.599071Z digest=sha256:df05506cc9c43d8f72bbab00151aaf5d85da9a40fe368c42df1633abf29a7184

Observation 7f616ca9-4ad1-4b58-9440-50b8e78a14f2 · outbound

This paper cites Common informa- tion based markov perfect equilibria for stochastic games with asymmetric information: Finite games.

Provable Partially Observable Reinforcement Learning with Privileged Information Common informa- tion based markov perfect equilibria for stochastic games with asymmetric information: Finite games

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.313480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.602837Z digest=sha256:a87ca03762e5322a4ca456d9f7b8f98ca399c0099bcc1bbabfe6ba9ef52cf19a

Observation af3b9a8a-b593-4bc1-812a-06b5c13054bd · outbound

This paper cites Decentralized stochastic con- trol with partial history sharing: A common information approach.

Provable Partially Observable Reinforcement Learning with Privileged Information Decentralized stochastic con- trol with partial history sharing: A common information approach

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.301288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.606535Z digest=sha256:e7c1d18e010156ccdbfe8c008b335020d774bb14a1b82b30e50b9cdc88d37f12

Observation cd4d198c-2a00-4d7f-a17b-ade838f12842 · outbound

This paper cites Sample-e fficient reinforcement learning of par- tially observable Markov games.

Provable Partially Observable Reinforcement Learning with Privileged Information Sample-e fficient reinforcement learning of par- tially observable Markov games

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.289167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.610519Z digest=sha256:f6b9818b756277f943af922ed4bb2b0b105b0bfa8db2390678681b9c7718fa4d

Observation 4df6750a-8ad8-4153-98ff-fc5ce1da62c0 · outbound

This paper cites When Can We Learn General-Sum Markov Games with a Large Number of Players Sample-Efficiently?.

Provable Partially Observable Reinforcement Learning with Privileged Information When Can We Learn General-Sum Markov Games with a Large Number of Players Sample-Efficiently?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.614294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.614294Z digest=sha256:48b52a80916dc13bd8d1a63889ec5754b49d5102591515fa173ee264ce8be867

Observation 6ec2666b-1ae8-49a4-bb75-441e8eb21da6 · outbound

This paper cites A sharp analysis of model-based reinforce- ment learning with self-play.

Provable Partially Observable Reinforcement Learning with Privileged Information A sharp analysis of model-based reinforce- ment learning with self-play

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.275849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.618444Z digest=sha256:48efb510c08e356e228d338521986262ffa2f50aa87cea9608217cbc9a320dd7

Observation 38b72e48-4179-4bd1-ad2b-17174936a734 · outbound

This paper cites V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL.

Provable Partially Observable Reinforcement Learning with Privileged Information V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.622093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.622093Z digest=sha256:c492ef1eba579b0102a7d46a078ec1ff99e6b8c02331243b0cb96f6aad9915d7

Observation b2b1acf5-a996-4570-8c66-3f65afcea8f3 · outbound

This paper cites Algorithmic game theory.

Provable Partially Observable Reinforcement Learning with Privileged Information Algorithmic game theory

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.263260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.626139Z digest=sha256:06312bbfc2f738e5f5bd3e7ab4265dc155889555a41b5af6e0b2bf4ca414d66b

Observation 413df1b0-1e47-4182-8beb-ead88dbc369f · outbound

This paper cites Kakade, and Yishay Mansour.

Provable Partially Observable Reinforcement Learning with Privileged Information Kakade, and Yishay Mansour

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.251251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.629852Z digest=sha256:92a5c4dc3c7e06d9adb48325b6baf5ebe3075a689863564844ad0acc231f6e5b

Observation 4a866159-e575-4a45-9ed7-802c1031900e · outbound

This paper cites Common information based markov perfect equilibria for linear-gaussian games with asymmetric information.

Provable Partially Observable Reinforcement Learning with Privileged Information Common information based markov perfect equilibria for linear-gaussian games with asymmetric information

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.238544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.633736Z digest=sha256:1f7898573f59b63f8637309f0f9840cc55b35ccd40732ad183fdc63593c3e63a

Observation cedc9884-654f-48df-a2a5-a4435ef11285 · outbound

This paper cites Poste- rior sampling for competitive rl: Function approximation and partial observation.

Provable Partially Observable Reinforcement Learning with Privileged Information Poste- rior sampling for competitive rl: Function approximation and partial observation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.226186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.637547Z digest=sha256:e9e688ee42d37cd01ead1c8e0d69f6d2ae99a3e740ed5bc524e0f2de83bdc7cb

Observation fa8a4d42-0c85-45f8-aa69-ebeebc845115 · outbound

This paper cites Learning to communicate with deep multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning to communicate with deep multi-agent reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.213731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.641729Z digest=sha256:c06ccdc4f171436bd2076de45a5f28f670cafb69a60e9ae08f2f9c80dcd546cb

Observation 2231e7d5-4078-439a-8c17-b7b62cfe3929 · outbound

This paper cites Computationally efficient pac rl in pomdps with latent determinism and conditional embeddings.

Provable Partially Observable Reinforcement Learning with Privileged Information Computationally efficient pac rl in pomdps with latent determinism and conditional embeddings

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.201238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.645575Z digest=sha256:d5e49be734c759ae7d10b6e13fee81d39cb0b01dbe0b623d9dc6a8d44c359d91

Observation e14c806f-8f65-42cf-9715-667326470196 · outbound

This paper cites Actor-critic algorithms.

Provable Partially Observable Reinforcement Learning with Privileged Information Actor-critic algorithms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.186967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.649339Z digest=sha256:546fd2a7017edf2f82d4ad451174f9f9743b180f07a94322adc563080d42d9ca

Observation 82d5fc1e-41a3-4ca6-8134-d17fa4c918e2 · outbound

This paper cites Provably e fficient exploration in policy optimization.

Provable Partially Observable Reinforcement Learning with Privileged Information Provably e fficient exploration in policy optimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.173029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.653116Z digest=sha256:4e0c8177cef4c94ce58ad16076994d5c677e09b8b283364241c2bea6fe8c6b23

Observation 31f7914a-92bf-4466-a3a6-50437b3e9e20 · outbound

This paper cites Optimistic policy optimization with bandit feedback.

Provable Partially Observable Reinforcement Learning with Privileged Information Optimistic policy optimization with bandit feedback

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.159754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.657114Z digest=sha256:53c6e791fbec7b391b96f1a5a64b847ae8a985442a0252528cb2106beb4d7537

Observation c73b445d-dab2-4326-8fb5-54831e4d0d0f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Provable Partially Observable Reinforcement Learning with Privileged Information Proximal Policy Optimization Algorithms

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.660877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.660877Z digest=sha256:51c6f76a1b660ed671e91e3a9ba8b617d508a7d4ef6484cf709e2ac52780fa00

Observation 4b321497-10f6-4430-a6b4-a709545abb65 · outbound

This paper cites A natural policy gradient.

Provable Partially Observable Reinforcement Learning with Privileged Information A natural policy gradient

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.146247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.664830Z digest=sha256:7ac059b0ab6eb56735052125ee6c5362d94c2e4e8ea8a694c6fa6179659ad38a

Observation 3c907966-e501-4a25-b882-e415f34b9484 · outbound

This paper cites Optimality and approxi- mation with policy gradient methods in Markov decision processes.

Provable Partially Observable Reinforcement Learning with Privileged Information Optimality and approxi- mation with policy gradient methods in Markov decision processes

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.131943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.668539Z digest=sha256:0902fec7e962344b49266ef9ce8511661e7c97899c8e1f48919d9bb391656f52

Observation 292cda01-7a8f-4ac7-a0f0-cfb0b2c95d91 · outbound

This paper cites Information state embedding in partially observable cooperative multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Information state embedding in partially observable cooperative multi-agent reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.119280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.672466Z digest=sha256:eca263bc6f143e89941e834245b6904c853b9ed74a6ba796bb707b604af36c1a

Observation c95e5aa1-8890-4093-9d8c-740e3d16bc7e · outbound

This paper cites Approximate infor- mation state for approximate planning and reinforcement learning in partially observed sys- tems.

Provable Partially Observable Reinforcement Learning with Privileged Information Approximate infor- mation state for approximate planning and reinforcement learning in partially observed sys- tems

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.106613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.676320Z digest=sha256:c122050f4685bd9567e2aa9213518157a554f51b7da3cc2ca37cfb17629a01aa

Observation c9bc6d5d-eefc-4818-be95-83c0487b7a49 · outbound

This paper cites Stochastic games with one step delay sharing information pattern with application to power control.

Provable Partially Observable Reinforcement Learning with Privileged Information Stochastic games with one step delay sharing information pattern with application to power control

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.094199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.680053Z digest=sha256:53a82f3ae118d559c687b303dfebecf5a50407937b392f7ef734e7cd3c4bbffb

Observation 9e3bb496-8db5-4244-830e-f8d3f22febbf · outbound

This paper cites A mea- surement study of internet delay asymmetry.

Provable Partially Observable Reinforcement Learning with Privileged Information A mea- surement study of internet delay asymmetry

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.081640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.683976Z digest=sha256:5e6a9936ecca93e50ade2d010f39114cf9f4f674fb8be11001a894238c17ecbb

Observation 6d2cb1c1-a240-4700-9c95-5033cda07d6b · outbound

This paper cites Repeated games with incomplete information.

Provable Partially Observable Reinforcement Learning with Privileged Information Repeated games with incomplete information

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.687754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.687754Z digest=sha256:039c2d3e82be86dafc84cd8531a4b6848363b086ad363809a579ec3c2da6f3f2

Observation 254b8c83-d1ab-4357-a907-d2f66ae616c0 · outbound

This paper cites Information theory: From coding to learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Information theory: From coding to learning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.060933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.691607Z digest=sha256:a1ac943a53d9a43fdaa98204241a44b2b70f483e0cd8ba91ca2b10d85a0eb217

Observation 67936a39-66f6-4d20-b77b-b4d56d62b14f · outbound

This paper cites On Value Functions and the Agent-Environment Boundary.

Provable Partially Observable Reinforcement Learning with Privileged Information On Value Functions and the Agent-Environment Boundary

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.695389Z digest=sha256:64835cd5d07373d29382f607f16160246277a57825a36de04bc222336408a271

Observation 6c07b385-7de8-4e98-8727-ce3a58b94cd4 · outbound

This paper cites A short note on learning discrete distributions.

Provable Partially Observable Reinforcement Learning with Privileged Information A short note on learning discrete distributions

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.699272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.699272Z digest=sha256:e6cc35f21b6fed67d42f06b79aee8dd687eb3c2ba062a3a1129e5dd00bcd1ec3

Observation 2580b214-3fde-481e-aa8e-d04beaee22e3 · outbound

This paper cites A characteri- zation of multiclass learnability, 2022.

Provable Partially Observable Reinforcement Learning with Privileged Information A characteri- zation of multiclass learnability, 2022

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.048694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.703594Z digest=sha256:d155a509cdc0daba0d043a5906a0eb99ed37390fcd48fa4c1e093f5f6b712890

Observation cffa59e6-104b-47c8-99c6-0c4b76a227a0 · outbound

This paper cites A characteriza- tion of multiclass learnability.

Provable Partially Observable Reinforcement Learning with Privileged Information A characteriza- tion of multiclass learnability

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.036051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.707289Z digest=sha256:874e9daa1762a61f240ba2bace6c7d146acba29e03661cd286f0baaf505e0715

Observation 88707e5b-c738-43f6-a625-0f666337a471 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Approximately optimal approximate reinforcement learning

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.023529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.711055Z digest=sha256:d60c0d104c6de08ecd51758607f774ea35495df2230a28005943f98ad699def5

Observation db7590f0-0b5f-43a8-81bb-880ea14462cf · outbound

This paper cites Convex optimization: Algorithms and complexity.

Provable Partially Observable Reinforcement Learning with Privileged Information Convex optimization: Algorithms and complexity

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.010854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.714868Z digest=sha256:e59ec459b0f28b0f5120884b06f1306be0eca5c4ac73c70c84ec9fa05c196afa

Observation c550fa58-6a96-4029-9677-f99375c1814a · outbound

This paper cites Tighter problem-dependent regret bounds in reinforce- ment learning without domain knowledge using value function bounds.

Provable Partially Observable Reinforcement Learning with Privileged Information Tighter problem-dependent regret bounds in reinforce- ment learning without domain knowledge using value function bounds

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.997919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.718762Z digest=sha256:f25735d813ee8c55aefa566d57d924b4148b3bc89b5ba5eb1e5e2a75323c94c3

Observation b84e87ff-29a4-42be-843a-e0b026fa63cb · outbound

This paper cites Reward-free exploration for reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Reward-free exploration for reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.984880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.722481Z digest=sha256:5f135c2003dd422cee6b07442c27dc5aad2fb1b4ed0bc7ce2f789db010444a54

Observation a754894a-2951-40e7-909b-970f2e2adf3b · outbound

This paper cites Is Q-learning provably efficient? In Advances in Neural Information Processing Systems, pages 4863–4873, 2018.

Provable Partially Observable Reinforcement Learning with Privileged Information Is Q-learning provably efficient? In Advances in Neural Information Processing Systems, pages 4863–4873, 2018

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.971151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.726414Z digest=sha256:958b9494888705e572d7fb99011881b836134ed5b7ebe05aea8421bed87cdc0b

Observation c3b66ef2-81f7-40af-ad94-b46acf6a31e8 · outbound

This paper cites No-regret learning in convex games.

Provable Partially Observable Reinforcement Learning with Privileged Information No-regret learning in convex games

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.958482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.730158Z digest=sha256:5c462ba07640b2e2b3692157f90cd7f6dda4c684cb236c50d8b9556fc367400a

Observation 2822478c-8fe3-4201-995a-0bdb4b02f0e7 · outbound

This paper cites No-regret learning in bayesian games.

Provable Partially Observable Reinforcement Learning with Privileged Information No-regret learning in bayesian games

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.945880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.734176Z digest=sha256:f10d229d4fcec525cf5d238b090bf3f767a7159ef3dd3633fa5df52f9fb0328b

Observation 1b60dc6e-6f1f-43bb-bd5c-1f1a7e5e29ab · outbound

This paper cites Bayes correlated equilibria, no-regret dynamics in Bayesian games, and the price of anarchy.

Provable Partially Observable Reinforcement Learning with Privileged Information Bayes correlated equilibria, no-regret dynamics in Bayesian games, and the price of anarchy

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.737884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.737884Z digest=sha256:89e21a03b499cd6f4552aca0173a9338dec3ae0f38ffa007352c403d84959d5f

Observation c3edf6af-5c31-413e-a19b-9db9d4b3320c · outbound

This paper cites By pluggingLt−1(π) into Equation (C.1), with simple algebric manipulations, we prove that: πt h(·|τh)∝πt−1 h (·|τh)exp ηEsh∼bh(τh) h Qt−1 h (τh,sh,·) i.

Provable Partially Observable Reinforcement Learning with Privileged Information By pluggingLt−1(π) into Equation (C.1), with simple algebric manipulations, we prove that: πt h(·|τh)∝πt−1 h (·|τh)exp ηEsh∼bh(τh) h Qt−1 h (τh,sh,·) i

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.933014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.742983Z digest=sha256:e0b25aeb8752a2635754ff0f7056e4164ef1d9fe868fb62e94fea4c2aba2f77a

Observation 1032c0d1-fd32-4015-a79e-d617f6fb4d71 · outbound

This paper cites bV (mi⋄πk i )⊙πk −i,G i,h+1 (ch+1) # ,H−h + 1 ) , where the last step is by inductive hypothesis. Now note that for anysh,ph,ah, we have bk−1 h (sh,ah) +Eoh+1∼bJk−1 h (·|sh,ah).

Provable Partially Observable Reinforcement Learning with Privileged Information bV (mi⋄πk i )⊙πk −i,G i,h+1 (ch+1) # ,H−h + 1 ) , where the last step is by inductive hypothesis. Now note that for anysh,ph,ah, we have bk−1 h (sh,ah) +Eoh+1∼bJk−1 h (·|sh,ah)

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T05:00:57.920214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:00:57.747243Z digest=sha256:284b59a55ea51e928c00a37c8b29602f72baf8f61cb4748c0c23d3b3fa530df7

Pith citing papers

No inbound Pith citation observations are available.