Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot Reinforcement Learning Under Partial Observability

As of 19 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 1 inbound Pith citation observation for arXiv:2506.15446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15446 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:37:29.933624Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T20:37:36.030165Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:37:45.261549Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47aec00e-6502-43af-8802-3139c8a455aa · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Zero-Shot Reinforcement Learning Under Partial Observability Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.511834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.511834Z digest=sha256:8cfa9e1781f8497402f920bc1de562c13243b8ed55e6e2bc28b687d8be1c708d

Observation 399ac21f-4282-415f-829e-b21d30473c3c · outbound

This paper cites Hindsight experience replay.

Zero-Shot Reinforcement Learning Under Partial Observability Hindsight experience replay

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.517302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.517302Z digest=sha256:85ffd6a36a7112f156bc47486530cb3c6f481309364d252dd2351564fa081e2b

Observation 3c5e6fef-941e-41fc-9407-1dfda27694b3 · outbound

This paper cites Rudder: Return decomposition for delayed rewards.

Zero-Shot Reinforcement Learning Under Partial Observability Rudder: Return decomposition for delayed rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.522097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.522097Z digest=sha256:1d45ccbdb756ff021fda1224577e991d109c1df7e9c51b8be61a060013e4d38b

Observation 0e4fac02-d175-476a-ab54-53c6b373f85b · outbound

This paper cites Optimal control of markov processes with incomplete state information.

Zero-Shot Reinforcement Learning Under Partial Observability Optimal control of markov processes with incomplete state information

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.526925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.526925Z digest=sha256:a947dce585022e8875d179603006be0bea2a155d34a22cfcbac35d0b79eec7f9

Observation 45d2ae62-49ed-481a-a99a-32d37107d400 · outbound

This paper cites Layer Normalization.

Zero-Shot Reinforcement Learning Under Partial Observability Layer Normalization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.531705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.531705Z digest=sha256:31c9c584909c58c3d313544db3b880727404e53db1a7b5541628e16d3dcc2786

Observation 5958cc51-6da9-45ea-a146-5e56ed537992 · outbound

This paper cites Reinforcement learning with long short-term memory.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement learning with long short-term memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.536697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.536697Z digest=sha256:152f03a33430977f88c29f8e8973d9f0d50d1a93f6aaf499ad6751612abc9432

Observation 5d7598e5-dc2f-4475-b4e7-6791b449ef44 · outbound

This paper cites Augmented world models facilitate zero-shot dynamics generalization from a single offline environment.

Zero-Shot Reinforcement Learning Under Partial Observability Augmented world models facilitate zero-shot dynamics generalization from a single offline environment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.541631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.541631Z digest=sha256:548ae810345c592f6558a1c33952b92e843b1f237b615a55a8096c508e92ebf3

Observation df418c7a-74d2-42fc-90e0-d4f9e9370ef7 · outbound

This paper cites Successor features for transfer in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Successor features for transfer in reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.545990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.545990Z digest=sha256:a4697b5a950eccfe9370fe31fbe69f0d2fab1d43841484f88cd57bf89b360b29

Observation 807bbc2b-591e-476e-91d1-f08bbd0a2476 · outbound

This paper cites Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint.

Zero-Shot Reinforcement Learning Under Partial Observability Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.550432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.550432Z digest=sha256:e7cb6262abdf267fbee7a8907d0781232f023e81ea53383df260964ccfe35e42

Observation 6f7f19c0-edcf-4ebc-8739-6b0b931a0ca5 · outbound

This paper cites Universal Successor Features Approximators.

Zero-Shot Reinforcement Learning Under Partial Observability Universal Successor Features Approximators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.555410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.555410Z digest=sha256:d3dd602c981f30c04ee5405ee74cbee60e233f0d4020487ff3ea470562d8aa2d

Observation 0f20e5ef-211b-4ca3-8052-946d5f61aa52 · outbound

This paper cites Language models are few-shot learners.

Zero-Shot Reinforcement Learning Under Partial Observability Language models are few-shot learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.560123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.560123Z digest=sha256:1d29ce321ff0ed31caf3eec8f45125da61f0ed34a8bec22b5a8a73651862e1c3

Observation 7a51e508-d1cb-472c-b9eb-e962433713ee · outbound

This paper cites Acting optimally in partially observable stochastic domains.

Zero-Shot Reinforcement Learning Under Partial Observability Acting optimally in partially observable stochastic domains

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.564489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.564489Z digest=sha256:0d43df2dce4fdd41f8e14f1fd3ce5757787d17e7a7c7eeba78e083b84d07df3a

Observation 18a9c2a8-f72a-47df-b564-5b3464cbbf7d · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Zero-Shot Reinforcement Learning Under Partial Observability Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.568951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.568951Z digest=sha256:e020d574cb2a0a843cd2f0c1068333881326c9b8b288ea24f3d31ef248906dc6

Observation 7cf032f4-5977-4097-86b2-29188d804dcc · outbound

This paper cites Quantifying generalization in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Quantifying generalization in reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.573854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.573854Z digest=sha256:4cfd2cd0acc7c3129285dc909587bc3bdc4d9defd5a54207f318dd8637f40580

Observation 71f1e26c-54b4-4b34-a55c-77caca9acb67 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Zero-Shot Reinforcement Learning Under Partial Observability Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.578537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.578537Z digest=sha256:06b623759b8fbdd8dc04985a60cb2430f4067c0f7061985b4d8c2244fa943f94

Observation 018f2ccd-ea2b-45d1-9a4a-544035a6d135 · outbound

This paper cites Improving generalization for temporal difference learning: The successor representation.

Zero-Shot Reinforcement Learning Under Partial Observability Improving generalization for temporal difference learning: The successor representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.583044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.583044Z digest=sha256:591a4db89057607eea8310d46adc349ab7e1b3a291e90f1947e004839ab6c066

Observation 917a2b4b-1c8a-4327-8789-f3a3e511695e · outbound

This paper cites Facing off world model backbones: Rnns, transformers, and s4.

Zero-Shot Reinforcement Learning Under Partial Observability Facing off world model backbones: Rnns, transformers, and s4

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.587523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.587523Z digest=sha256:37ec61930475b529713ee4b95852cdaf38457487df56788dc50f5a53b8477463

Observation 5fd46fdc-f1da-44db-a719-a2b49aabea3c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Zero-Shot Reinforcement Learning Under Partial Observability An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.592608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.592608Z digest=sha256:7e2b652e5f6807914e33c1f3371dc9f07fecfbf0db209e49b34b2d34b19a246e

Observation c3a71902-fbe1-4608-abd0-7f0320dfcb80 · outbound

This paper cites Finding structure in time.

Zero-Shot Reinforcement Learning Under Partial Observability Finding structure in time

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.597239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.597239Z digest=sha256:784c44bb475facf270898cf704d67630b033372acc9f1b3f39d5f4dee6f10c63

Observation 0b34e4d8-bf29-442c-80a1-6bf691cf607f · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Contrastive learning as goal-conditioned reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.601741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.601741Z digest=sha256:dc11a5a91e1a11c8df7449199a596eae4b0a0b1e33dd8186ca9f86527c94e3d7

Observation 73fc2854-b51f-4b46-831d-6332dc046698 · outbound

This paper cites Generalization and Regularization in DQN.

Zero-Shot Reinforcement Learning Under Partial Observability Generalization and Regularization in DQN

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.607617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.607617Z digest=sha256:05a661a8a1c145b095693fc3e5246f174e232ee285b283d2881b368725332118

Observation 619b6e4f-3fe6-4971-9ecb-e1c500ca43f0 · outbound

This paper cites Hyperbolic Discounting and Learning over Multiple Horizons.

Zero-Shot Reinforcement Learning Under Partial Observability Hyperbolic Discounting and Learning over Multiple Horizons

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.612462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.612462Z digest=sha256:656da92aa525eb1e966c698c4451cd29ed9b6ed9c3a0ec309fd91183b3271fc9

Observation fde3a438-0752-4baf-9aa8-c1f66459ba30 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability A minimalist approach to offline reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.617136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.617136Z digest=sha256:c63a5aebe3421c0a1d159162bf4bd71f5ec0320329d8c1e789a0c07f2edac980

Observation 2f4615da-27d2-4969-8652-9ce997cf7745 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Zero-Shot Reinforcement Learning Under Partial Observability Off-policy deep reinforcement learning without exploration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.621574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.621574Z digest=sha256:d70f6af3459b906bdfc4d2eae61169acb9c73556fae30a0f197b5b4201be234a

Observation 868ea262-eba7-4a19-a0e2-33fb4a5319d3 · outbound

This paper cites Amago: Scalable in-context reinforcement learning for adaptive agents.

Zero-Shot Reinforcement Learning Under Partial Observability Amago: Scalable in-context reinforcement learning for adaptive agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:31.024078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.625814Z digest=sha256:5f6d10086fe84ba27b3e02f09e6925db2f18d66a7e1d1ec9c1c394fabe72b965

Observation 19a6fd89-7eff-4f64-8b40-6bb77964c9f3 · outbound

This paper cites Amago-2: Breaking the multi-task barrier in meta-reinforcement learning with transformers.

Zero-Shot Reinforcement Learning Under Partial Observability Amago-2: Breaking the multi-task barrier in meta-reinforcement learning with transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:31.007717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.631547Z digest=sha256:ed43bf35fcd3724c399cce7cb553e53ecc1a6906b9d1a80b243b9534bb57a944

Observation 35a9efee-c19a-4029-87b3-3a21b3c93456 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Zero-Shot Reinforcement Learning Under Partial Observability Efficiently Modeling Long Sequences with Structured State Spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.636036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.636036Z digest=sha256:db41ee20f8daff982ce24661c640c6e9d62a88d806f0ab2f04ec462f7442470a

Observation 1a2375f2-8882-4cbb-b369-39a29f030b87 · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

Zero-Shot Reinforcement Learning Under Partial Observability On the parameterization and initialization of diagonal state space models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.640688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.640688Z digest=sha256:447f1e460d1da063c035689c4a8d4fa73a61c4bead88dfa21de3afbac26a658e

Observation 4e96b70a-c51c-4dac-b4a4-430d444f4b44 · outbound

This paper cites World Models.

Zero-Shot Reinforcement Learning Under Partial Observability World Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.644594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.644594Z digest=sha256:a627f9e6f41b9b21276d88a18ae9fe4f3439fb73b18dae2007ed8436cc965bbd

Observation cec1eb11-adb9-4954-b9fa-cd406450e963 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Zero-Shot Reinforcement Learning Under Partial Observability Dream to Control: Learning Behaviors by Latent Imagination

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.649175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.649175Z digest=sha256:2864dc36b1409e81414250767900e7b44eea788c1e8471af93b052ef4b01aa5f

Observation a4494390-9afd-4372-8eb9-a8cce066cecb · outbound

This paper cites Learning latent dynamics for planning from pixels.

Zero-Shot Reinforcement Learning Under Partial Observability Learning latent dynamics for planning from pixels

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.982829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.653652Z digest=sha256:56dea4725854d5fd76bcd4c970dd5b6ee5613d25732173c694603b9f8971aa7e

Observation e7f9fcb4-037f-4854-97dd-5005adc42819 · outbound

This paper cites Mastering Atari with Discrete World Models.

Zero-Shot Reinforcement Learning Under Partial Observability Mastering Atari with Discrete World Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.657706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.657706Z digest=sha256:f15dbcaff219688844b20a5926a423ebb9f767ed28c7d827d68b7e4b80ec8dce

Observation 526d1092-4fa5-44ec-9586-d9fc7d2e24ab · outbound

This paper cites Mastering Diverse Domains through World Models.

Zero-Shot Reinforcement Learning Under Partial Observability Mastering Diverse Domains through World Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.662121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.662121Z digest=sha256:d3f9ad8cfac5e45c949e65c7d1996712b13d541d65473b96ab4b729acc11fdc7

Observation a6803eb5-81af-46ef-a626-26547d7140fe · outbound

This paper cites Contextual markov decision processes, 2015.

Zero-Shot Reinforcement Learning Under Partial Observability Contextual markov decision processes, 2015

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.666961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.666961Z digest=sha256:f82f71882a6e6ae1135f0af1cc5fdaaa0af17361d4452f00db1e967c24644ebb

Observation 958323c3-a7a9-4fc2-8b65-0c48db1a8552 · outbound

This paper cites Array programming with numpy.

Zero-Shot Reinforcement Learning Under Partial Observability Array programming with numpy

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.671261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.671261Z digest=sha256:0a7567789ab153c6793035250a5098e8424bd014fddab12c3c081e3029f74f6d

Observation 6e5cdc20-671d-4396-9ac3-3f6f6fe3e72c · outbound

This paper cites Deep recurrent q-learning for partially observable mdps.

Zero-Shot Reinforcement Learning Under Partial Observability Deep recurrent q-learning for partially observable mdps

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.675694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.675694Z digest=sha256:484f7cfd7595f09134fd36ac6088aedddf9edf185d48a8a42241d80d454e954a

Observation dffbaf06-fc6c-47b2-8540-f7d978c67333 · outbound

This paper cites Memory-based control with recurrent neural networks.

Zero-Shot Reinforcement Learning Under Partial Observability Memory-based control with recurrent neural networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.680133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.680133Z digest=sha256:d8e9fc12ddbba978206068a2e752fd701b8d1d35dfe786f990fa161faa0e1fdb

Observation 20cc5211-af52-4287-a5dc-add1e0fdd9f5 · outbound

This paper cites Long short-term memory.

Zero-Shot Reinforcement Learning Under Partial Observability Long short-term memory

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.939846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.684679Z digest=sha256:1929fca794d739652662dd87cc6fc00d8e2aed586077ac17d124423e644da5c1

Observation b4a9fb5e-b0bf-410b-beb2-096d3fc161c5 · outbound

This paper cites Matplotlib: A 2d graphics environment.

Zero-Shot Reinforcement Learning Under Partial Observability Matplotlib: A 2d graphics environment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.688888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.688888Z digest=sha256:099571db11119ef9e741b51c470d6f15f557ae516d7ccf20b1c98be7eb9dc0d3

Observation d63be35e-9830-451e-a8fa-8f4d1fd35ff8 · outbound

This paper cites an unresolved cited work.

Zero-Shot Reinforcement Learning Under Partial Observability Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:37:30.914737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.693207Z digest=sha256:98fdd398a149b917d0c3b13feb817f365a0aec866227de43735f0d0f41a7b08a

Observation 81c701c8-fac6-4b18-a580-8f718be115a5 · outbound

This paper cites Monotonic robust policy optimization with model discrepancy.

Zero-Shot Reinforcement Learning Under Partial Observability Monotonic robust policy optimization with model discrepancy

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.899413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.697415Z digest=sha256:9417f1d6a1b8cfc04f52a5d35f62f8f35fd523e68969a75f3e318b289f8dc5c3

Observation fc5dd288-84eb-4f1f-b2be-b8867e5896dc · outbound

This paper cites Planning and acting in partially observable stochastic domains.

Zero-Shot Reinforcement Learning Under Partial Observability Planning and acting in partially observable stochastic domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.701983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.701983Z digest=sha256:3ce3296a5338c06f3974f48ff16df7a3d3673e4ba0f6d5a4a72754fd43e73eaf

Observation c6177a93-3c17-4a27-bb5f-5158d13e8ce0 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Morel: Model-based offline reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.873696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.706390Z digest=sha256:b824c7ee5f26be9e25a7fedfa935a66371d9b0f0066997bfbcccbc15299cf96f

Observation e2565ba5-0f27-43d9-bbc7-5caaa321bd93 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Zero-Shot Reinforcement Learning Under Partial Observability Adam: A Method for Stochastic Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.710772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.710772Z digest=sha256:63162422b5ad0fd4c64bdba3caaf022836ac963d95c7373294685ff9fbe4bcb1

Observation 316fda0c-1f91-4d51-92dd-1a97e660c5b5 · outbound

This paper cites Actor-critic algorithms.

Zero-Shot Reinforcement Learning Under Partial Observability Actor-critic algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.715191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.715191Z digest=sha256:4f414c46b54f525bf132c82db8172365f3e4bc5ccfae3bb88205600badbe6ce9

Observation 0aab0761-f32c-49a6-bd90-a5ccccec3cfa · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Zero-Shot Reinforcement Learning Under Partial Observability Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.849393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.719538Z digest=sha256:c04120967a2f217bcd968b2b89193096dea6567851b1e8a9302d1f698b5510fd

Observation da367ce7-9459-4610-a7ca-ceacf3493761 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Zero-Shot Reinforcement Learning Under Partial Observability Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.834686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.723906Z digest=sha256:ce4d595d4ce9bb3e4eb601214c4b90f6cd7d09bf0ce6a131cefff7763fd9a4cd

Observation 5cfd831a-980d-4ee2-9d6a-d92b81f6c87e · outbound

This paper cites Conservative Q-Learning for Offline Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Conservative Q-Learning for Offline Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.728341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.728341Z digest=sha256:9ad80debad26aa15cd378c47527a944f8a95f2366c79fa02b9fa7afba7bb88ec

Observation 0067d450-c1b3-4c8f-b663-c6fc30b9e3a1 · outbound

This paper cites Batch reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Batch reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.818157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.732981Z digest=sha256:89e7035083aa00ec2936f554ffe567ed400a646f45504f70df68d4d4a90586de

Observation 5e06f878-525f-472e-adad-f795481a4fb4 · outbound

This paper cites Context-aware dynamics model for generalization in model-based reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Context-aware dynamics model for generalization in model-based reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.802726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.737738Z digest=sha256:e7928713944c1c35c6bd282eeebe5f58aced6b1997cd59acf5670623dcc11e39

Observation e4664fa5-2a55-44d1-ae18-dab3ef0786dc · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Zero-Shot Reinforcement Learning Under Partial Observability Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.742221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.742221Z digest=sha256:774a8c28d1f945c472b7bd6ec7b2607e84bfc6fd5941d8828064d67333522d2c

Observation a9df9c00-9c55-42e5-beb8-3fba3fee2770 · outbound

This paper cites Off-Policy Policy Gradient with State Distribution Correction.

Zero-Shot Reinforcement Learning Under Partial Observability Off-Policy Policy Gradient with State Distribution Correction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.746768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.746768Z digest=sha256:2acc15dedc638c210f03d892903f73f5e7ed123fb6c9ca68321d8349b3c104b6

Observation c1e78f30-178c-4e9e-ba34-f1a035390933 · outbound

This paper cites Structured state space models for in-context reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Structured state space models for in-context reinforcement learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.787481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.751518Z digest=sha256:0d457fe7f7362915b13e64879c08ea880510b0e76a4be776bb6843fe5dd49910

Observation fa96089b-dc21-49ce-8d75-5eb1e52daf8a · outbound

This paper cites How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression.

Zero-Shot Reinforcement Learning Under Partial Observability How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.755654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.755654Z digest=sha256:790f1a4120a42c4cbd35a43ec112f2d6cd2711d09c36cd0dec9b9b18b025c5b6

Observation 08750780-abcb-422c-beff-5bb4857cb255 · outbound

This paper cites Robust Reinforcement Learning for Continuous Control with Model Misspecification.

Zero-Shot Reinforcement Learning Under Partial Observability Robust Reinforcement Learning for Continuous Control with Model Misspecification

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.760121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.760121Z digest=sha256:4119334a74316c61e02a2c9f425f59a2ea540291ee81fdec2fd0def8b4f082e8

Observation b6325162-4e18-4290-9751-450323571248 · outbound

This paper cites pandas: a foundational python library for data analysis and statistics.

Zero-Shot Reinforcement Learning Under Partial Observability pandas: a foundational python library for data analysis and statistics

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.772487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.764802Z digest=sha256:7c0e7379c8d2e2baaa6d10f5701a3c30672d97305b353d09d2bb1883105c7d8a

Observation 3cd82194-0b93-42c2-bb5a-11296011faa1 · outbound

This paper cites Memory-based deep reinforcement learning for pomdps.

Zero-Shot Reinforcement Learning Under Partial Observability Memory-based deep reinforcement learning for pomdps

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.756148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.769141Z digest=sha256:3743e0d7978bc0d95499f35ea414e61361d1d8b3c3c7c83be896f05e03ef842c

Observation 4acd8763-474e-426d-b984-32f648861b0a · outbound

This paper cites Steps toward artificial intelligence.

Zero-Shot Reinforcement Learning Under Partial Observability Steps toward artificial intelligence

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.773832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.773832Z digest=sha256:b647c7512af4b13a83e5d80be3b256916d781198220121324f6f2882bf0b09c8

Observation a60f4992-812f-4624-9136-81254a918629 · outbound

This paper cites Human-level control through deep reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Human-level control through deep reinforcement learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.778130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.778130Z digest=sha256:b2814a1d47448300a2d582202862386e93689b31f722079739779ba40b03069c

Observation 323129d7-14c5-4449-bc98-43eccd4e5638 · outbound

This paper cites POPGym: Benchmarking Partially Observable Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability POPGym: Benchmarking Partially Observable Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.782550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.782550Z digest=sha256:bcd2bb8c6efcc04ad43cd7db8fa5beeb7cf798bdb7860901e3f0a91abf43c6dc

Observation 49a2e96b-4457-41b0-942b-d83da35f6a7f · outbound

This paper cites Robust reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Robust reinforcement learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.719577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.787224Z digest=sha256:ffc44df4d9c4d3c56ac21af5bd0a141a39739fd1d47b896fd39089c44c118ac1

Observation c7d8530c-a575-4793-9688-dcc3157acc13 · outbound

This paper cites Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs.

Zero-Shot Reinforcement Learning Under Partial Observability Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.792082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.792082Z digest=sha256:aac1e66a2055db341322cb76bd7e900762ce7d84b2ae36cff2700dcdbcfe390c

Observation 4b262d7c-34f8-46c8-a5b2-fae1dd7b0b27 · outbound

This paper cites Robust control of markov decision processes with uncertain transition matrices.

Zero-Shot Reinforcement Learning Under Partial Observability Robust control of markov decision processes with uncertain transition matrices

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.796830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.796830Z digest=sha256:3b5d28e1c776a25ed0d637c286f2c2bda345b7b2889012e5edfa8ccde555c5ca

Observation 68186d0b-af03-4957-908b-1541299b95e9 · outbound

This paper cites Assessing Generalization in Deep Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Assessing Generalization in Deep Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.801239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.801239Z digest=sha256:c5cb04d1e04e794d476377ab6ca52c764460f30e884c91731b303ff1fd46f587

Observation 60a194a3-7aac-4a71-81bc-1a8db8f6cf3d · outbound

This paper cites Stabilizing transformers for reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Stabilizing transformers for reinforcement learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.805746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.805746Z digest=sha256:71413f3fa82a37a578d293ef9afb40cac48d5faca6a7c2f31b25d429578d3358

Observation e71ba0c2-2617-423e-a59a-b25b1ca40278 · outbound

This paper cites Hiql: Offline goal-conditioned rl with latent states as actions.

Zero-Shot Reinforcement Learning Under Partial Observability Hiql: Offline goal-conditioned rl with latent states as actions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.683624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.810869Z digest=sha256:914adf4ee13e7ac5f1eb484ab276baf5fa6915bbcdd0b2662ffb09047df433c0

Observation 8e683216-8398-4ecf-b387-9b5f5c20c041 · outbound

This paper cites Foundation policies with hilbert representations.

Zero-Shot Reinforcement Learning Under Partial Observability Foundation policies with hilbert representations

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.668577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.815410Z digest=sha256:45e023d1a44953ed9e2531f43153824c6dc82f5d3d13698845a9fdc331b072ac

Observation 511fde79-6680-448b-840a-3c87033fd7c6 · outbound

This paper cites Automatic differentiation in pytorch.

Zero-Shot Reinforcement Learning Under Partial Observability Automatic differentiation in pytorch

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.819694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.819694Z digest=sha256:51bce14098fc9f15227aaa59a881c36adaa862687995e03a7bac0aa02b70c2a4

Observation b5e4fd5d-7023-4e3b-bee2-b7e4f92dca37 · outbound

This paper cites Fast imitation via behavior foundation models.

Zero-Shot Reinforcement Learning Under Partial Observability Fast imitation via behavior foundation models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.644796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.824782Z digest=sha256:7c4b5c3c6e772c0e4844e6182ecb4129045eef6e81805bdfe210101c7f5eba00

Observation 55abfbbe-c184-4cb9-8ee9-edf199a201e8 · outbound

This paper cites Automatic Data Augmentation for Generalization in Deep Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Automatic Data Augmentation for Generalization in Deep Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.829176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.829176Z digest=sha256:e5c18b5dfc2b1f6cb12726f31038651e80534ebf4cd8524560e899515555bbbe

Observation 2227ed27-a674-4b16-a5a0-b4056ed2ff5c · outbound

This paper cites EPOpt: Learning Robust Neural Network Policies Using Model Ensembles.

Zero-Shot Reinforcement Learning Under Partial Observability EPOpt: Learning Robust Neural Network Policies Using Model Ensembles

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.833746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.833746Z digest=sha256:043fb5e8f16597dfc137d87c4dee482147afe5e95ae37102b29bb7d2a08f58d8

Observation c2c8a653-fcf0-4190-8bda-e8640f52f4f0 · outbound

This paper cites Synthetic Returns for Long-Term Credit Assignment.

Zero-Shot Reinforcement Learning Under Partial Observability Synthetic Returns for Long-Term Credit Assignment

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.838312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.838312Z digest=sha256:faa151ca246b6d6b4f1ecbeacd8e0a708a8e9fa7ae445ef925e4ee3997602d1f

Observation 09b5a9d9-404b-451b-8af9-e16407be0569 · outbound

This paper cites Reward-Free Curricula for Training Robust World Models.

Zero-Shot Reinforcement Learning Under Partial Observability Reward-Free Curricula for Training Robust World Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.842674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.842674Z digest=sha256:0de869aeaa96ca0c3963fde8b82dbf8b7a1185a9c88b16d8d8fd402c41ae6b7d

Observation 81b78502-8f0f-4f5a-8a45-426ca853d253 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Zero-Shot Reinforcement Learning Under Partial Observability High-resolution image synthesis with latent diffusion models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.847403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.847403Z digest=sha256:a0372eaacfb1edba0e4ac056fe836a48bfd32c9cced4b7e03c70c79a93e22242

Observation 4936e410-ef66-4b55-92fb-5520789b7837 · outbound

This paper cites Python: a programming language for software integration and development.

Zero-Shot Reinforcement Learning Under Partial Observability Python: a programming language for software integration and development

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.620707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.851790Z digest=sha256:ed0d48de9865810f32bbaaa6addab1c6fd0693f9a5b26dd0a175bc290222cf7a

Observation 376148d4-a25c-4f5f-a7c9-9358faec5095 · outbound

This paper cites Universal value function approximators.

Zero-Shot Reinforcement Learning Under Partial Observability Universal value function approximators

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.605450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.856128Z digest=sha256:71e21ffff7db28c3b8bd510185d5d1b506f41cc31cff07833d89dd75d862f849

Observation 38fcec82-510a-4a84-9c06-39722501636c · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.860626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.860626Z digest=sha256:ac9be83cbfef3bd4e3de5b835461483180ea71e848e2b50204f97cfe6c9f93f9

Observation b01114b7-071b-4cea-91f1-65115677c446 · outbound

This paper cites Reinforcement learning in markovian and non-markovian environments.

Zero-Shot Reinforcement Learning Under Partial Observability Reinforcement learning in markovian and non-markovian environments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.590304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.865474Z digest=sha256:5138c7811bfd71d9d661e42c86181e23d5ce52dab092c30b3ead5e54bb5dd92b

Observation 1ecd03db-429e-4f0b-95c7-02500b6b485e · outbound

This paper cites Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.574754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.869702Z digest=sha256:3c6c67493c2aec1aadee5d72659eae47a93ff1de0903974c5d7013f602309f6d

Observation 407b2b38-e276-4752-86a2-bfd964238e72 · outbound

This paper cites Temporal credit assignment in reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Temporal credit assignment in reinforcement learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.873991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.873991Z digest=sha256:7dd96c717925dc983304aafba037eef7070300b4c2c798a897c194647bd3652b

Observation 2a2f8887-3798-40b0-8beb-df1ae25b0f29 · outbound

This paper cites DeepMind Control Suite.

Zero-Shot Reinforcement Learning Under Partial Observability DeepMind Control Suite

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.878368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.878368Z digest=sha256:adb1dd3402e1afcac87f84a3621d5d1f1efea92adebde66705cca7317905bc93

Observation 6d4e7c52-71e6-4868-b5af-3b9062571e5d · outbound

This paper cites Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.883154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.883154Z digest=sha256:9b23e3524b187e8edf60aea5c66c3a4e9d8bc7a97c144d45cf5b3275c783c0af

Observation 47b145d8-d437-4e1b-9037-cdcc5700b3f9 · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world.

Zero-Shot Reinforcement Learning Under Partial Observability Domain randomization for transferring deep neural networks from simulation to the real world

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.888157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.888157Z digest=sha256:11d9a693994d19b2d68acc80b63aea10e17c0a6d22e1bdc8adc3337cab2476ac

Observation a2beca31-fc93-478e-9096-a643b663d7be · outbound

This paper cites Mujoco: A physics engine for model-based control.

Zero-Shot Reinforcement Learning Under Partial Observability Mujoco: A physics engine for model-based control

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.892663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.892663Z digest=sha256:62e3c6ff500a04acf5391d5734a2052e10ff3f530ca5415af82dc01e19726c9c

Observation 966fed24-1e51-4619-a0e3-1cf3138a950e · outbound

This paper cites Learning one representation to optimize all rewards.

Zero-Shot Reinforcement Learning Under Partial Observability Learning one representation to optimize all rewards

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.530779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.897013Z digest=sha256:a09a9b1ad71fbe9875e30dda2a587421adc54fb2ba699b576e8711d903e056b3

Observation e5f34fca-cc7e-4d0e-acd8-be2ac9b95955 · outbound

This paper cites Does zero-shot reinforcement learning exist? In The Eleventh International Conference on Learning Representations, 2023.

Zero-Shot Reinforcement Learning Under Partial Observability Does zero-shot reinforcement learning exist? In The Eleventh International Conference on Learning Representations, 2023

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.515552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.901534Z digest=sha256:df4505d0a4a0c9079c8f576f81ad47fc31017aae82837f6472c238fe68508939

Observation de40c939-25da-42d0-8f94-1af2e11fe73f · outbound

This paper cites Attention is all you need.

Zero-Shot Reinforcement Learning Under Partial Observability Attention is all you need

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.905823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.905823Z digest=sha256:a92226c6e2f4c93a08b58f7284e15cfbab09913ca8782a4af14f0c14123abdca

Observation a13be115-0c55-47d0-ac2e-63fdc5df8801 · outbound

This paper cites Active perception and reinforcement learning.

Zero-Shot Reinforcement Learning Under Partial Observability Active perception and reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.490669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.910566Z digest=sha256:e8660d748c6f6035535a24650841d544756cc4dcc8ce6a860e8e3ff28b382228

Observation e5be8a5a-1179-4702-9d08-d29702676e7a · outbound

This paper cites Policy gradient critics.

Zero-Shot Reinforcement Learning Under Partial Observability Policy gradient critics

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.475027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.915226Z digest=sha256:52d261070ff6e2aab72e40a490d7b3a720c0af4005e36b02e56f730e7b308d58

Observation 2c43ebbf-cdcb-438d-a90c-89b5e2d43f8b · outbound

This paper cites Meta-gradient reinforcement learning with an objective discovered online.

Zero-Shot Reinforcement Learning Under Partial Observability Meta-gradient reinforcement learning with an objective discovered online

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.460421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.919944Z digest=sha256:f2f493e0ba7b0eec279d1efbfbeee87b3ad9b76d71d099fe03775b2c5aa75eee

Observation 4e439a89-65e0-422c-8c60-62163ed0f744 · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Zero-Shot Reinforcement Learning Under Partial Observability Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.924574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.924574Z digest=sha256:aae044cfb2aabf9f760be34c1f2a01b34f8dcd776a1572c22e77035fbae97817

Observation 765afd43-5b4c-48f8-8968-bbf63bc2da0a · outbound

This paper cites Learning deep neural network policies with continuous memory states.

Zero-Shot Reinforcement Learning Under Partial Observability Learning deep neural network policies with continuous memory states

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:37:30.445202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:37:29.929217Z digest=sha256:182a43f69e8fc68f2795e4d4b4bc31ce6bda9c979856d36d3fc3308feddb4ac0

Observation 66ed7fcb-7b1b-4894-b12f-6225b0543ca4 · outbound

This paper cites write newline.

Zero-Shot Reinforcement Learning Under Partial Observability write newline

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T19:37:29.933624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:37:29.933624Z digest=sha256:449ddb1a0c04d9eac9df8e7875b8761fd8bb0d341b0b1b5b3f07bc79063359a0

Pith citing papers

Observation 50626c59-d028-4721-9d66-d21892112808 · inbound

When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited cites this paper.

When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited Zero-Shot Reinforcement Learning Under Partial Observability

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:37:45.263201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T20:37:36.030165Z digest=sha256:7e0785636b5e9b4267a86e601fc77fc4c178d26dae97da6bd443633c8b2a5984