Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot Reinforcement Learning Under Partial Observability

As of 17 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 1 inbound Pith citation observation for arXiv:2506.15446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15446 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:37:29.933624Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T20:37:36.030165Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:37:45.261549Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47aec00e-6502-43af-8802-3139c8a455aa · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Zero-Shot Reinforcement Learning Under Partial Observability Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.511834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.511834Z digest=sha256:b81f92d50c3113c5cfab47d116565128178c0769eaf73247e302f5f697e55edb

Observation 399ac21f-4282-415f-829e-b21d30473c3c · outbound

This paper cites Hindsight experience replay.

Zero-Shot Reinforcement Learning Under Partial Observability Hindsight experience replay

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.517302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.517302Z digest=sha256:dceae75335386fbd276c42fca8ed10ba93ef7cf8cbd7df819835c264a5e1d0b8

Observation 3c5e6fef-941e-41fc-9407-1dfda27694b3 · outbound

This paper cites Rudder: Return decomposition for delayed rewards.

Zero-Shot Reinforcement Learning Under Partial Observability Rudder: Return decomposition for delayed rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.522097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.522097Z digest=sha256:976ad2d4c191e5c2f10df2be4f530ebf00b58d7f7516422de78827c7d0e725a0

Observation 0e4fac02-d175-476a-ab54-53c6b373f85b · outbound

This paper cites Optimal control of markov processes with incomplete state information.

Zero-Shot Reinforcement Learning Under Partial Observability Optimal control of markov processes with incomplete state information

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.526925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.526925Z digest=sha256:bb87e6d5410b6d8e661d19d9d735483652706b924c4945fbc404b3c10bd2873d

Observation 45d2ae62-49ed-481a-a99a-32d37107d400 · outbound

This paper cites Layer Normalization.

Zero-Shot Reinforcement Learning Under Partial Observability Layer Normalization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.531705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.531705Z digest=sha256:2c691c47d6eaa12f927e71b865450adf81fefca07b4f821848de1138bf533cb9

Observation 5958cc51-6da9-45ea-a146-5e56ed537992 · outbound

This paper cites Reinforcement learning with long short-term memory.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement learning with long short-term memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.536697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.536697Z digest=sha256:6d66c817445d0274bb1d701f42cec243164714e7ce29aa3f0483453a37dfdb3f

Observation 5d7598e5-dc2f-4475-b4e7-6791b449ef44 · outbound

This paper cites Augmented world models facilitate zero-shot dynamics generalization from a single offline environment.

Zero-Shot Reinforcement Learning Under Partial Observability Augmented world models facilitate zero-shot dynamics generalization from a single offline environment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.541631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.541631Z digest=sha256:e673f8e2c0992ce7a3080303122370e18c615d9062f85278c599f43972c24ef1

Observation df418c7a-74d2-42fc-90e0-d4f9e9370ef7 · outbound

This paper cites Successor features for transfer in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Successor features for transfer in reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.545990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.545990Z digest=sha256:6c15fc586a212605bb6cfe1e5d85b1380356e6516aa87bd7cbf7fff2852ec070

Observation 807bbc2b-591e-476e-91d1-f08bbd0a2476 · outbound

This paper cites Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint.

Zero-Shot Reinforcement Learning Under Partial Observability Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.550432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.550432Z digest=sha256:b47406b2e070e08b96976ed4e69f5910a538d940041633357a9e0ff0908bd6e5

Observation 6f7f19c0-edcf-4ebc-8739-6b0b931a0ca5 · outbound

This paper cites Universal Successor Features Approximators.

Zero-Shot Reinforcement Learning Under Partial Observability Universal Successor Features Approximators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.555410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.555410Z digest=sha256:ddf2a60adeaa17ee551fe79fc0e75a3da214eb8bf0decdd2313ccb7bc55d1e9b

Observation 0f20e5ef-211b-4ca3-8052-946d5f61aa52 · outbound

This paper cites Language models are few-shot learners.

Zero-Shot Reinforcement Learning Under Partial Observability Language models are few-shot learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.560123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.560123Z digest=sha256:a1d3c982ad1c7a3902da145512a68ab0af536c27f763c89132fb19cca9a8b222

Observation 7a51e508-d1cb-472c-b9eb-e962433713ee · outbound

This paper cites Acting optimally in partially observable stochastic domains.

Zero-Shot Reinforcement Learning Under Partial Observability Acting optimally in partially observable stochastic domains

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.564489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.564489Z digest=sha256:fa84c2851c4d2c3226f5388d9fc8ad06df5179501fe820a960e5c2c42b364cbf

Observation 18a9c2a8-f72a-47df-b564-5b3464cbbf7d · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Zero-Shot Reinforcement Learning Under Partial Observability Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.568951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.568951Z digest=sha256:800c052f53305c8721d52679678c9eb7f9d87ccd5f7b5d511a3e45c2536de908

Observation 7cf032f4-5977-4097-86b2-29188d804dcc · outbound

This paper cites Quantifying generalization in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Quantifying generalization in reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.573854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.573854Z digest=sha256:515a2b6641817246d72ebc6ee559039efb7730d55f4fe5750b06721268cd460a

Observation 71f1e26c-54b4-4b34-a55c-77caca9acb67 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Zero-Shot Reinforcement Learning Under Partial Observability Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.578537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.578537Z digest=sha256:a4275e184342319025ac3d678488ae1f628b0b4e81d63e80a33ba44fe7c56aba

Observation 018f2ccd-ea2b-45d1-9a4a-544035a6d135 · outbound

This paper cites Improving generalization for temporal difference learning: The successor representation.

Zero-Shot Reinforcement Learning Under Partial Observability Improving generalization for temporal difference learning: The successor representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.583044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.583044Z digest=sha256:dbec25e48fb6063fd2f8785f7ba6456ed36a4e0cba4d2ab0d419c574e6bf81ed

Observation 917a2b4b-1c8a-4327-8789-f3a3e511695e · outbound

This paper cites Facing off world model backbones: Rnns, transformers, and s4.

Zero-Shot Reinforcement Learning Under Partial Observability Facing off world model backbones: Rnns, transformers, and s4

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.587523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.587523Z digest=sha256:e64f883b3145045825f7fcf0ac8c62112366c7697d353a6b588b5a30f0e5ee0a

Observation 5fd46fdc-f1da-44db-a719-a2b49aabea3c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Zero-Shot Reinforcement Learning Under Partial Observability An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.592608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.592608Z digest=sha256:797c3eb57c697f8d2f50cd6b187ef6f14323f5aafb7f60102a7d4824dff1bb18

Observation c3a71902-fbe1-4608-abd0-7f0320dfcb80 · outbound

This paper cites Finding structure in time.

Zero-Shot Reinforcement Learning Under Partial Observability Finding structure in time

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.597239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.597239Z digest=sha256:ea53d9947f4d59d0c94fdada74f16b910270922a71dc986c1fcc419f29153b0f

Observation 0b34e4d8-bf29-442c-80a1-6bf691cf607f · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Contrastive learning as goal-conditioned reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.601741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.601741Z digest=sha256:eb33f809a48f7264a99163272b610e3e7de0cca3cb39e005ec723456aa208b67

Observation 73fc2854-b51f-4b46-831d-6332dc046698 · outbound

This paper cites Generalization and Regularization in DQN.

Zero-Shot Reinforcement Learning Under Partial Observability Generalization and Regularization in DQN

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.607617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.607617Z digest=sha256:2f6eddc8f335025216e4c1789230b9b62de53207f69ade19fa10be7348f0d3b0

Observation 619b6e4f-3fe6-4971-9ecb-e1c500ca43f0 · outbound

This paper cites Hyperbolic Discounting and Learning over Multiple Horizons.

Zero-Shot Reinforcement Learning Under Partial Observability Hyperbolic Discounting and Learning over Multiple Horizons

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.612462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.612462Z digest=sha256:b21763999c3e38354bdc4012fb29f68c263e24ce8db50567938a60badfbc2fb3

Observation fde3a438-0752-4baf-9aa8-c1f66459ba30 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability A minimalist approach to offline reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.617136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.617136Z digest=sha256:d4254bd82fa23ecb3bf5d6289378e0c864f89bd126d43f6da2873e11f374305c

Observation 2f4615da-27d2-4969-8652-9ce997cf7745 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Zero-Shot Reinforcement Learning Under Partial Observability Off-policy deep reinforcement learning without exploration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.621574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.621574Z digest=sha256:f0e5825cb3c6c9c3aec4d009b01913c374e620ce8bfd4e9891b99f25e4916630

Observation 868ea262-eba7-4a19-a0e2-33fb4a5319d3 · outbound

This paper cites Amago: Scalable in-context reinforcement learning for adaptive agents.

Zero-Shot Reinforcement Learning Under Partial Observability Amago: Scalable in-context reinforcement learning for adaptive agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:31.024078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.625814Z digest=sha256:c826b823d0b46eda8c823b5cbd5e29c03bcd9cfdc5910ae4d664509710ac43a5

Observation 19a6fd89-7eff-4f64-8b40-6bb77964c9f3 · outbound

This paper cites Amago-2: Breaking the multi-task barrier in meta-reinforcement learning with transformers.

Zero-Shot Reinforcement Learning Under Partial Observability Amago-2: Breaking the multi-task barrier in meta-reinforcement learning with transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:31.007717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.631547Z digest=sha256:1527686af5f0b74bd7d9ac367f9354244b063acb8608ec25edf9cf8f159304d7

Observation 35a9efee-c19a-4029-87b3-3a21b3c93456 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Zero-Shot Reinforcement Learning Under Partial Observability Efficiently Modeling Long Sequences with Structured State Spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.636036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.636036Z digest=sha256:e1dd9bd2ed4e00ee20014813f0867705f6ca9e8f70b535f175d623dcea12de16

Observation 1a2375f2-8882-4cbb-b369-39a29f030b87 · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

Zero-Shot Reinforcement Learning Under Partial Observability On the parameterization and initialization of diagonal state space models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.640688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.640688Z digest=sha256:2bf5a4cd7086c003ebc90954e48d8f73fa8bf8698e2476e3b9582c5fba1fde90

Observation 4e96b70a-c51c-4dac-b4a4-430d444f4b44 · outbound

This paper cites World Models.

Zero-Shot Reinforcement Learning Under Partial Observability World Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.644594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.644594Z digest=sha256:52187f6f312f652178d8e54357720192f673eed04ac66e873f501fa9d2027183

Observation cec1eb11-adb9-4954-b9fa-cd406450e963 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Zero-Shot Reinforcement Learning Under Partial Observability Dream to Control: Learning Behaviors by Latent Imagination

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.649175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.649175Z digest=sha256:7ecc7b6ae6391cc125b30d073fe1102e20eded00dbf4060abf471f78d7680a64

Observation a4494390-9afd-4372-8eb9-a8cce066cecb · outbound

This paper cites Learning latent dynamics for planning from pixels.

Zero-Shot Reinforcement Learning Under Partial Observability Learning latent dynamics for planning from pixels

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.982829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.653652Z digest=sha256:cedd592b46bee6167c0fc24aa17169fbd38035d7f3a94eba91d3bf19402c71d8

Observation e7f9fcb4-037f-4854-97dd-5005adc42819 · outbound

This paper cites Mastering Atari with Discrete World Models.

Zero-Shot Reinforcement Learning Under Partial Observability Mastering Atari with Discrete World Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.657706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.657706Z digest=sha256:d9f9e1fcb34d1f78d89ec60bc695b990a8af42c9b3841025429ecb7d4e137b4b

Observation 526d1092-4fa5-44ec-9586-d9fc7d2e24ab · outbound

This paper cites Mastering Diverse Domains through World Models.

Zero-Shot Reinforcement Learning Under Partial Observability Mastering Diverse Domains through World Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.662121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.662121Z digest=sha256:799aca523659e24ca3258c6212f480f35e26151109b6993de924452afc0a76df

Observation a6803eb5-81af-46ef-a626-26547d7140fe · outbound

This paper cites Contextual markov decision processes, 2015.

Zero-Shot Reinforcement Learning Under Partial Observability Contextual markov decision processes, 2015

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.666961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.666961Z digest=sha256:50a0d852e8befecac812a33ea57f7db96497443a11b29343ebded1f4bc54f217

Observation 958323c3-a7a9-4fc2-8b65-0c48db1a8552 · outbound

This paper cites Array programming with numpy.

Zero-Shot Reinforcement Learning Under Partial Observability Array programming with numpy

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.671261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.671261Z digest=sha256:71630ea268734d51a0d55bdece0e4c2ba3d7a6d628808e3ea14d1d89aa4ef237

Observation 6e5cdc20-671d-4396-9ac3-3f6f6fe3e72c · outbound

This paper cites Deep recurrent q-learning for partially observable mdps.

Zero-Shot Reinforcement Learning Under Partial Observability Deep recurrent q-learning for partially observable mdps

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.675694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.675694Z digest=sha256:89465830cb524b852ba8cf5cc20d81f1278bdcffd67f99587164184ab1921d52

Observation dffbaf06-fc6c-47b2-8540-f7d978c67333 · outbound

This paper cites Memory-based control with recurrent neural networks.

Zero-Shot Reinforcement Learning Under Partial Observability Memory-based control with recurrent neural networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.680133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.680133Z digest=sha256:d546f45b610240ea6f10e4bf68b075deb9776b3f21bfc8ff21636da7681ee048

Observation 20cc5211-af52-4287-a5dc-add1e0fdd9f5 · outbound

This paper cites Long short-term memory.

Zero-Shot Reinforcement Learning Under Partial Observability Long short-term memory

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.939846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.684679Z digest=sha256:891b340876d9e4612e4a89d9b0d97dce05cd63f85e3eb9e1685469df5faf209f

Observation b4a9fb5e-b0bf-410b-beb2-096d3fc161c5 · outbound

This paper cites Matplotlib: A 2d graphics environment.

Zero-Shot Reinforcement Learning Under Partial Observability Matplotlib: A 2d graphics environment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.688888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.688888Z digest=sha256:09594e31f73edf832fd56a214f2c8fe832ca08ba6ced8c28a3938ca24d8194f2

Observation d63be35e-9830-451e-a8fa-8f4d1fd35ff8 · outbound

This paper cites an unresolved cited work.

Zero-Shot Reinforcement Learning Under Partial Observability Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:37:30.914737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.693207Z digest=sha256:e899aae81c4fe6334837a6acc04e5820420ec2add51a043cde951dd88c9441e4

Observation 81c701c8-fac6-4b18-a580-8f718be115a5 · outbound

This paper cites Monotonic robust policy optimization with model discrepancy.

Zero-Shot Reinforcement Learning Under Partial Observability Monotonic robust policy optimization with model discrepancy

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.899413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.697415Z digest=sha256:7ed13ceb25b3da5a7634c085c58d04080f1ddb321aeb482be1bf585577b031b7

Observation fc5dd288-84eb-4f1f-b2be-b8867e5896dc · outbound

This paper cites Planning and acting in partially observable stochastic domains.

Zero-Shot Reinforcement Learning Under Partial Observability Planning and acting in partially observable stochastic domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.701983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.701983Z digest=sha256:5a76f5b77003f1f109b3491499a304bc1272149e664a17c2c97b61025082445f

Observation c6177a93-3c17-4a27-bb5f-5158d13e8ce0 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Morel: Model-based offline reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.873696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.706390Z digest=sha256:7db084356cd68e10a2eb9eb78a9c738e8548ba93050f13404c0154b9250cde9a

Observation e2565ba5-0f27-43d9-bbc7-5caaa321bd93 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Zero-Shot Reinforcement Learning Under Partial Observability Adam: A Method for Stochastic Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.710772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.710772Z digest=sha256:622b48badf8f5ef0887b23cc99e347f61edf2341de70263a258077225c17822e

Observation 316fda0c-1f91-4d51-92dd-1a97e660c5b5 · outbound

This paper cites Actor-critic algorithms.

Zero-Shot Reinforcement Learning Under Partial Observability Actor-critic algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.715191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.715191Z digest=sha256:461dd0cbe96d04de3a97880909f9f34d24d497756caaf9b6db8745b3494d0fb0

Observation 0aab0761-f32c-49a6-bd90-a5ccccec3cfa · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Zero-Shot Reinforcement Learning Under Partial Observability Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.849393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.719538Z digest=sha256:fcdf398dbb16aefa765e6b3f4fdffb35744146efd5dcdd751dc5cf79d6266ba0

Observation da367ce7-9459-4610-a7ca-ceacf3493761 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Zero-Shot Reinforcement Learning Under Partial Observability Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.834686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.723906Z digest=sha256:56b143c045e6d3abaae8f09c4c007aa9f94780a2bfc6162884e46d434930836c

Observation 5cfd831a-980d-4ee2-9d6a-d92b81f6c87e · outbound

This paper cites Conservative Q-Learning for Offline Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Conservative Q-Learning for Offline Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.728341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.728341Z digest=sha256:bae987a6d15cb213bb7aa4847d1294cb108282d0e030b203f0d803c750b03647

Observation 0067d450-c1b3-4c8f-b663-c6fc30b9e3a1 · outbound

This paper cites Batch reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Batch reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.818157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.732981Z digest=sha256:a397dd29c616ba73d27666a46c7491c59c042b37f10f472c43fae4bf35fd277c

Observation 5e06f878-525f-472e-adad-f795481a4fb4 · outbound

This paper cites Context-aware dynamics model for generalization in model-based reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Context-aware dynamics model for generalization in model-based reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.802726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.737738Z digest=sha256:33e29eb302d5c49cb69d4920ebab0ffa81b6e3e59eac1e95ff742e1317d6c2a2

Observation e4664fa5-2a55-44d1-ae18-dab3ef0786dc · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Zero-Shot Reinforcement Learning Under Partial Observability Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.742221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.742221Z digest=sha256:7a1ccfa226bc6b6a2168e3a5ed24f5de417114e5c6266755f27c4361c102a6d5

Observation a9df9c00-9c55-42e5-beb8-3fba3fee2770 · outbound

This paper cites Off-Policy Policy Gradient with State Distribution Correction.

Zero-Shot Reinforcement Learning Under Partial Observability Off-Policy Policy Gradient with State Distribution Correction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.746768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.746768Z digest=sha256:dd07d8f684d774a76b8c02ba118e6ae66b60ca9439b8595dc01d68c8251bc25b

Observation c1e78f30-178c-4e9e-ba34-f1a035390933 · outbound

This paper cites Structured state space models for in-context reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Structured state space models for in-context reinforcement learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.787481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.751518Z digest=sha256:0d4fc656a2fc6b427c4cefdda77a1e729bcf39acf7c637f134756e16a3e1ffa8

Observation fa96089b-dc21-49ce-8d75-5eb1e52daf8a · outbound

This paper cites How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression.

Zero-Shot Reinforcement Learning Under Partial Observability How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.755654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.755654Z digest=sha256:2a4ac24a3b4540436378dcb40eafe24e69fa78f3b94aeb3f83aba992a5425937

Observation 08750780-abcb-422c-beff-5bb4857cb255 · outbound

This paper cites Robust Reinforcement Learning for Continuous Control with Model Misspecification.

Zero-Shot Reinforcement Learning Under Partial Observability Robust Reinforcement Learning for Continuous Control with Model Misspecification

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.760121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.760121Z digest=sha256:c3a5bbbb2984a1b8014e72495fb923ada9c2fef6027dbb18d984caadc987a391

Observation b6325162-4e18-4290-9751-450323571248 · outbound

This paper cites pandas: a foundational python library for data analysis and statistics.

Zero-Shot Reinforcement Learning Under Partial Observability pandas: a foundational python library for data analysis and statistics

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.772487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.764802Z digest=sha256:761f55a16de4220409443140cd441e1c1e90f528fe363f5c1febb46dd20ca2f1

Observation 3cd82194-0b93-42c2-bb5a-11296011faa1 · outbound

This paper cites Memory-based deep reinforcement learning for pomdps.

Zero-Shot Reinforcement Learning Under Partial Observability Memory-based deep reinforcement learning for pomdps

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.756148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.769141Z digest=sha256:5da6c737d87cb65e4d3b653c05d788c5ba4219c78b395dc2f87e6f0cd4f11476

Observation 4acd8763-474e-426d-b984-32f648861b0a · outbound

This paper cites Steps toward artificial intelligence.

Zero-Shot Reinforcement Learning Under Partial Observability Steps toward artificial intelligence

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.773832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.773832Z digest=sha256:6be60aa980c864035da240fddc41e5b6924c6467ade5dc862b1745708253c070

Observation a60f4992-812f-4624-9136-81254a918629 · outbound

This paper cites Human-level control through deep reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Human-level control through deep reinforcement learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.778130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.778130Z digest=sha256:e4df81236dd63dc74532e0b1d1842627973f90831f07643b7a9646625f1747f5

Observation 323129d7-14c5-4449-bc98-43eccd4e5638 · outbound

This paper cites POPGym: Benchmarking Partially Observable Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability POPGym: Benchmarking Partially Observable Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.782550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.782550Z digest=sha256:7c73f81ce6d47628a53f2bd1bc11384678eca9b69725496f84f957c894ead796

Observation 49a2e96b-4457-41b0-942b-d83da35f6a7f · outbound

This paper cites Robust reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Robust reinforcement learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.719577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.787224Z digest=sha256:535ad941f4c79e4cd436744a1694639175c44a7250025a5f4ad3343ef6256963

Observation c7d8530c-a575-4793-9688-dcc3157acc13 · outbound

This paper cites Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs.

Zero-Shot Reinforcement Learning Under Partial Observability Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.792082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.792082Z digest=sha256:fa834dac66bc9fe12054d07572be28df4ac6a3ef984ba36be3aed855fa9b775c

Observation 4b262d7c-34f8-46c8-a5b2-fae1dd7b0b27 · outbound

This paper cites Robust control of markov decision processes with uncertain transition matrices.

Zero-Shot Reinforcement Learning Under Partial Observability Robust control of markov decision processes with uncertain transition matrices

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.796830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.796830Z digest=sha256:44df09ca8ee68084113fbeb57ec575f3baecb5964a88375e8553cf05d92aa863

Observation 68186d0b-af03-4957-908b-1541299b95e9 · outbound

This paper cites Assessing Generalization in Deep Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Assessing Generalization in Deep Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.801239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.801239Z digest=sha256:8f4afd325517556f9453da1577fe2cdcca1ccd56895131bc71d80a12bfd5c638

Observation 60a194a3-7aac-4a71-81bc-1a8db8f6cf3d · outbound

This paper cites Stabilizing transformers for reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Stabilizing transformers for reinforcement learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.805746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.805746Z digest=sha256:e58ebb34ac9f542a1d2ca5452c5ae4a0830414ccf462fc5ae04fc281b5955629

Observation e71ba0c2-2617-423e-a59a-b25b1ca40278 · outbound

This paper cites Hiql: Offline goal-conditioned rl with latent states as actions.

Zero-Shot Reinforcement Learning Under Partial Observability Hiql: Offline goal-conditioned rl with latent states as actions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.683624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.810869Z digest=sha256:6694e4d11459e5ffe47caa849c36aaea8357780e74721516bf4979c8dd76a224

Observation 8e683216-8398-4ecf-b387-9b5f5c20c041 · outbound

This paper cites Foundation policies with hilbert representations.

Zero-Shot Reinforcement Learning Under Partial Observability Foundation policies with hilbert representations

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.668577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.815410Z digest=sha256:30370fbf7b2690623af937dbb1bf255752f589dd1f91f1986704777b81d89883

Observation 511fde79-6680-448b-840a-3c87033fd7c6 · outbound

This paper cites Automatic differentiation in pytorch.

Zero-Shot Reinforcement Learning Under Partial Observability Automatic differentiation in pytorch

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.819694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.819694Z digest=sha256:6cd15b7f1675997e665ad2d476bf9d3af69521fb957969a7733477c6516a8404

Observation b5e4fd5d-7023-4e3b-bee2-b7e4f92dca37 · outbound

This paper cites Fast imitation via behavior foundation models.

Zero-Shot Reinforcement Learning Under Partial Observability Fast imitation via behavior foundation models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.644796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.824782Z digest=sha256:11a25000d71eb71e78918b5fc4744e928eae96b24e492ba94d0f2d3469698f24

Observation 55abfbbe-c184-4cb9-8ee9-edf199a201e8 · outbound

This paper cites Automatic Data Augmentation for Generalization in Deep Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Automatic Data Augmentation for Generalization in Deep Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.829176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.829176Z digest=sha256:4c8f1e492656457520be901e9bf5d05a68e65eb35c78d02e709cc5127ba9d4fe

Observation 2227ed27-a674-4b16-a5a0-b4056ed2ff5c · outbound

This paper cites EPOpt: Learning Robust Neural Network Policies Using Model Ensembles.

Zero-Shot Reinforcement Learning Under Partial Observability EPOpt: Learning Robust Neural Network Policies Using Model Ensembles

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.833746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.833746Z digest=sha256:3ef3f5ba6824381927e29398fbdcb165e870e25f12e428f424d855735b42ca19

Observation c2c8a653-fcf0-4190-8bda-e8640f52f4f0 · outbound

This paper cites Synthetic Returns for Long-Term Credit Assignment.

Zero-Shot Reinforcement Learning Under Partial Observability Synthetic Returns for Long-Term Credit Assignment

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.838312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.838312Z digest=sha256:d2a48ff3d9c9867533a34d28f7abab9d741d102b9ea1dd084ab78d4494bea266

Observation 09b5a9d9-404b-451b-8af9-e16407be0569 · outbound

This paper cites Reward-Free Curricula for Training Robust World Models.

Zero-Shot Reinforcement Learning Under Partial Observability Reward-Free Curricula for Training Robust World Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.842674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.842674Z digest=sha256:2c94144cec5b7b332a2da2b19b87038b75616bc6c0f289762cb206190ea068f4

Observation 81b78502-8f0f-4f5a-8a45-426ca853d253 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Zero-Shot Reinforcement Learning Under Partial Observability High-resolution image synthesis with latent diffusion models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.847403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.847403Z digest=sha256:09b1fc764c064cc8894a3cc3257627a2c1948fe873dbfc9873643cf7b2732eaf

Observation 4936e410-ef66-4b55-92fb-5520789b7837 · outbound

This paper cites Python: a programming language for software integration and development.

Zero-Shot Reinforcement Learning Under Partial Observability Python: a programming language for software integration and development

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.620707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.851790Z digest=sha256:1e979783d06b75434277fea8b24a44b9c4f2c72d1418489557f9cb61b490aa57

Observation 376148d4-a25c-4f5f-a7c9-9358faec5095 · outbound

This paper cites Universal value function approximators.

Zero-Shot Reinforcement Learning Under Partial Observability Universal value function approximators

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.605450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.856128Z digest=sha256:163d6b5938d4096a67648844c7c5e0bc5b9e14c43db7e8b748f541f9d8f5cb09

Observation 38fcec82-510a-4a84-9c06-39722501636c · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.860626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.860626Z digest=sha256:16253d7ecd6103991b2c4e612b74381e02c01c37e9e9ad5bcc26824f35c313f6

Observation b01114b7-071b-4cea-91f1-65115677c446 · outbound

This paper cites Reinforcement learning in markovian and non-markovian environments.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement learning in markovian and non-markovian environments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.590304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.865474Z digest=sha256:393627d25e32361d7c6b085ead88d0460babcdbe23e0e67d55e2016feb03189a

Observation 1ecd03db-429e-4f0b-95c7-02500b6b485e · outbound

This paper cites Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.574754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.869702Z digest=sha256:47dfc235e2831ee125b4c120a1f78f1b7e4aa7ac644d2e9d44fcfcb03ad3215e

Observation 407b2b38-e276-4752-86a2-bfd964238e72 · outbound

This paper cites Temporal credit assignment in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Temporal credit assignment in reinforcement learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.873991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.873991Z digest=sha256:c62dd1d46da32bb82a01d9d5e766881fd3e673e9b6c4db31b75a46bd88cdcf81

Observation 2a2f8887-3798-40b0-8beb-df1ae25b0f29 · outbound

This paper cites DeepMind Control Suite.

Zero-Shot Reinforcement Learning Under Partial Observability DeepMind Control Suite

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.878368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.878368Z digest=sha256:10a4c2fd286bfa6529cadd3c5554e51ade6f66327f139cfc869ddc3c58988665

Observation 6d4e7c52-71e6-4868-b5af-3b9062571e5d · outbound

This paper cites Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.883154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.883154Z digest=sha256:191d40056d15bf967ef719e78e0c90207a0c7449da3352f9c803f127322f1e89

Observation 47b145d8-d437-4e1b-9037-cdcc5700b3f9 · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world.

Zero-Shot Reinforcement Learning Under Partial Observability Domain randomization for transferring deep neural networks from simulation to the real world

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.888157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.888157Z digest=sha256:c6abbbfbd8a04ac85ec4cfdfcbcde84a21e5634a5f5d522fb8bfbff53ab475d1

Observation a2beca31-fc93-478e-9096-a643b663d7be · outbound

This paper cites Mujoco: A physics engine for model-based control.

Zero-Shot Reinforcement Learning Under Partial Observability Mujoco: A physics engine for model-based control

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.892663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.892663Z digest=sha256:6a436c1487221957e252f227e341faf311c684e0528f2cd889be0560f9e96fd3

Observation 966fed24-1e51-4619-a0e3-1cf3138a950e · outbound

This paper cites Learning one representation to optimize all rewards.

Zero-Shot Reinforcement Learning Under Partial Observability Learning one representation to optimize all rewards

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.530779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.897013Z digest=sha256:896f25dc232e469e38b77605780c4d2fc873c81f7f6aed94437d44857e6e1f30

Observation e5f34fca-cc7e-4d0e-acd8-be2ac9b95955 · outbound

This paper cites Does zero-shot reinforcement learning exist? In The Eleventh International Conference on Learning Representations, 2023.

Zero-Shot Reinforcement Learning Under Partial Observability Does zero-shot reinforcement learning exist? In The Eleventh International Conference on Learning Representations, 2023

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.515552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.901534Z digest=sha256:3107a712048f2e2ca0156643cc0581aacb781274fa3cf1b893484f5990f288b8

Observation de40c939-25da-42d0-8f94-1af2e11fe73f · outbound

This paper cites Attention is all you need.

Zero-Shot Reinforcement Learning Under Partial Observability Attention is all you need

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.905823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.905823Z digest=sha256:0b9669ae487b102bed310b5c938f3922f308db382b811ff2be7cda4ffdb239a3

Observation a13be115-0c55-47d0-ac2e-63fdc5df8801 · outbound

This paper cites Active perception and reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Active perception and reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.490669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.910566Z digest=sha256:0fa533b58a7f569bcb0f243a9236f033110d0b446ab34ee4665ce605cd1ce1cc

Observation e5be8a5a-1179-4702-9d08-d29702676e7a · outbound

This paper cites Policy gradient critics.

Zero-Shot Reinforcement Learning Under Partial Observability Policy gradient critics

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.475027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.915226Z digest=sha256:271acdfa664613923b2b38c1180a5810ae40b1fcdfcc1df8fd44d405e8c8d4d3

Observation 2c43ebbf-cdcb-438d-a90c-89b5e2d43f8b · outbound

This paper cites Meta-gradient reinforcement learning with an objective discovered online.

Zero-Shot Reinforcement Learning Under Partial Observability Meta-gradient reinforcement learning with an objective discovered online

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.460421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.919944Z digest=sha256:c705a70f586d13a845fbdbebfbd61631adb62f991215f243f56cd9ef7721cb3d

Observation 4e439a89-65e0-422c-8c60-62163ed0f744 · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.924574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.924574Z digest=sha256:6edb8821e540745c1f63cb155479f83a4a5f2b61263a66dce9efb0378ee7969c

Observation 765afd43-5b4c-48f8-8968-bbf63bc2da0a · outbound

This paper cites Learning deep neural network policies with continuous memory states.

Zero-Shot Reinforcement Learning Under Partial Observability Learning deep neural network policies with continuous memory states

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.445202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.929217Z digest=sha256:3fdb68286cdff8d945abaca5ec546f25e274b5c32b5f0545cf625ce97f8beefb

Observation 66ed7fcb-7b1b-4894-b12f-6225b0543ca4 · outbound

This paper cites write newline.

Zero-Shot Reinforcement Learning Under Partial Observability write newline

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.933624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.933624Z digest=sha256:edb14d943494cf3967b495a020d123ebf5407142bedd9bf248152d8882a32fff

Pith citing papers

Observation 50626c59-d028-4721-9d66-d21892112808 · inbound

When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited cites this paper.

When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited Zero-Shot Reinforcement Learning Under Partial Observability

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:37:45.263201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T20:37:36.030165Z digest=sha256:b85176966281f54eba3a5ead8701a0e4c7d8350d84b8d3e6705ecdde2311b40a