Pith. sign in

Paper Citation Record · LEDGER

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes

As of 20 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.11153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11153 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:22.868400Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1868bd64-3169-486a-ac63-7bef0f6ee600 · outbound

This paper cites Reinforcement learning in game industry—review, prospects and challenges,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement learning in game industry—review, prospects and challenges,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.706877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.596392Z digest=sha256:50e4da4654dc95c8982138e08357c72100d2bf470a797a94c684df0d960f5507

Observation 5b36d62b-6933-4947-ad82-0c273075a833 · outbound

This paper cites Artificial intelligence, machine learning and deep learning in advanced robotics, a review,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Artificial intelligence, machine learning and deep learning in advanced robotics, a review,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.601706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.601706Z digest=sha256:bff69157e76aafebdea2f14c84a0a3e8c2d34114d001bad3f4faa044c9509770

Observation afb15a91-ee65-49f1-81ff-92498e50dd95 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.606744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.606744Z digest=sha256:9b2c0e0d9848af5ccefee102ba793e6894bebf7510e29b28bfb2d426ca2022e9

Observation f0d16ee2-0918-4a41-9bc0-b48eeb8b0c50 · outbound

This paper cites Weakly coupled deep q-networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Weakly coupled deep q-networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.682572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.611371Z digest=sha256:3c3b1b8e262237a739c2d5327b8d1434db8e347e05a1f0b87b8ef2c4c792637d

Observation 2ca25150-c058-45a0-950c-fe3891c86abf · outbound

This paper cites Tuning apex dqn: A reinforcement learning based deep q-network algorithm,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Tuning apex dqn: A reinforcement learning based deep q-network algorithm,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.668685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.616440Z digest=sha256:1576cdf3f554b0dd96f2b04425e1fdbb88c57d56726c314a3c4db45f1f516cbc

Observation 0e111d3c-b700-45ee-8af1-df7aa143c6ae · outbound

This paper cites Safe reinforcement learning via shielding under partial observability,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Safe reinforcement learning via shielding under partial observability,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.654874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.620781Z digest=sha256:e030d98bce4bd1f4148f46609e547c229262e1d4b940d86cdb3617daadafa3f2

Observation c0c88722-5cad-4be5-be71-14d9dad8da2e · outbound

This paper cites Integrated task and motion planning for safe legged navigation in partially observable environments,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Integrated task and motion planning for safe legged navigation in partially observable environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.641601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.625634Z digest=sha256:4e40f6c39f5c498e0f878709811fdd4b8c667dcea913a861eb969aac05de2cf9

Observation 2a3f48f8-7f06-47b2-a650-a797f9fb2ba3 · outbound

This paper cites Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.629718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.629718Z digest=sha256:b8bd592dbec61f72b3c329fff18509e01612d05ebdeef874d77b108585485317

Observation 95640671-922e-4ad1-b63e-a3f8ee8426f4 · outbound

This paper cites A definition of continual reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes A definition of continual reinforcement learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.633844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.633844Z digest=sha256:6622cc70d4fe5021e7c07b8c38cb5e871ca9efb3edb062e858bca4d27fb42820

Observation 0d9b6466-e0f7-4ed8-a408-370001b6650b · outbound

This paper cites Deep reinforcement learning unleashing the power of ai in decision-making,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning unleashing the power of ai in decision-making,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.610494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.638178Z digest=sha256:cef18dc7b728204848b8dc50e5ce3f9efc6efd9fb2a82a9e1cf84f5deb89e139

Observation b554e27a-280c-45c4-89ba-5f6b8de6cda1 · outbound

This paper cites Exploration in deep reinforcement learning: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Exploration in deep reinforcement learning: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.642683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.642683Z digest=sha256:2ab5808a8c17571020584225d818c7567dd80a6539a0c848d9ab57ec774936b7

Observation ca5fec4b-57dc-47e4-bf84-fa36857fee63 · outbound

This paper cites Memory gym: Partially observable challenges to memory-based agents,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Memory gym: Partially observable challenges to memory-based agents,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.646843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.646843Z digest=sha256:dfda315c8c015e5b23f82a94884642dfaf6b34029b8bbdea34da73415a57a663

Observation 0657da13-58f0-4684-8e76-e35f23d2df8e · outbound

This paper cites Deep rein- forcement learning: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep rein- forcement learning: A survey,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.579178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.655825Z digest=sha256:b21d9f477a2fc20d2302da22fd6712f565ec8b86fbb1d94e69e40d15c6a2b901

Observation 5f247310-6da1-4525-b271-3e4b7100aa90 · outbound

This paper cites Partially Observable Markov Decision Processes (POMDPs) and Robotics.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Partially Observable Markov Decision Processes (POMDPs) and Robotics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.659923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.659923Z digest=sha256:fddab56c4db9f9f6d3969141b8910cf57f6be1d3b5006464271ecbd434dda909

Observation 179ec413-502f-4197-b9b0-e231b42dee0c · outbound

This paper cites Recurrent neural networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.565191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.664501Z digest=sha256:5fbc15133ad41166fa0caa7c0414f0ccb1ddd81280acf9d5c801137e16239a2f

Observation 5508d405-6ba2-44ff-a703-bd4ae0ce2130 · outbound

This paper cites Long short-term memory,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Long short-term memory,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.669081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.669081Z digest=sha256:8a7b98d8f110b1ce28341717cd780a6bb848dc4fc9b1e92ecf064abc782052a1

Observation 6c3462ee-db52-4081-a500-2e68d82f52db · outbound

This paper cites Gate-variants of gated recurrent unit (gru) neural networks,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Gate-variants of gated recurrent unit (gru) neural networks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.543519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.673326Z digest=sha256:1b234f0a22952442214c28d584c1364a05b8df2042aaadf1b51084ac5149c444

Observation 70c446ac-4926-4f2c-9ad1-327e144ea182 · outbound

This paper cites Recurrent prediction model for partially observable mdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent prediction model for partially observable mdps,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.529450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.677802Z digest=sha256:02d067e86527d0740fbe16f311b6beddf246c72e3ca5051b302ea8981e44ecaa

Observation ef5001db-3314-4c8c-b216-d89e3637eeda · outbound

This paper cites Visualizing transformers for nlp: a brief survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Visualizing transformers for nlp: a brief survey,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.514757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.682088Z digest=sha256:933ee802fb0b484195f4127bdc73131b90949ba6e2bb423dfd5aec33deb9053b

Observation c37841f9-3099-4022-80b6-56f081478351 · outbound

This paper cites Transformers in vision: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformers in vision: A survey,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.686804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.686804Z digest=sha256:3956c3300f4c2c30e568b9bb5a02d012663fef8b449acba67369d9f63d044aac

Observation cc20339b-6200-4c54-9e2d-9b4618c91561 · outbound

This paper cites Windows deep transformer q-networks: an extended variance reduction architecture for partially observable reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Windows deep transformer q-networks: an extended variance reduction architecture for partially observable reinforcement learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.490794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.691129Z digest=sha256:40acfb6f23c96e7724eb637ef4728fb89683a8300970f959f7594be50d69f83d

Observation c391dc60-aab2-49e3-89ab-0fe39343235a · outbound

This paper cites Deep Transformer Q-Networks for Partially Observable Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Transformer Q-Networks for Partially Observable Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.695510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.695510Z digest=sha256:3d840ce19145f610625890a54bfe2a7a6831680a438c6e0d3e5660f835942d40

Observation 3888cc66-5364-4345-93f6-9de049e8a60e · outbound

This paper cites Deep Recurrent Q-Learning for Partially Observable MDPs.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Recurrent Q-Learning for Partially Observable MDPs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.701126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.701126Z digest=sha256:0d6de76d82ecffea9044f615ada580ae53b1a9aa8bf0e336238d4b8bbd8912e9

Observation 72a74ba4-3b51-48b8-be89-0bedcb1171fc · outbound

This paper cites On improving deep reinforcement learning for pomdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On improving deep reinforcement learning for pomdps,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.476678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.705893Z digest=sha256:bc40ab504a228ae20a8a663885aa7e72c350623465f0f3e87b5de74820372707

Observation 46461b7f-00a1-4fad-b3cd-85734e91a628 · outbound

This paper cites Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.714601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.714601Z digest=sha256:b8164a2b6ed79f15aea40972e7acf69884c554c4ebeb02d8e85b96a78fbe5d73

Observation bf1657e5-c0f3-4a7c-b00d-852e039ba788 · outbound

This paper cites Deep reinforcement learning with bidirectional recurrent neural networks for dynamic spectrum access,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep reinforcement learning with bidirectional recurrent neural networks for dynamic spectrum access,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.463274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.719144Z digest=sha256:c80d4996d3fbc0224b840bde9dc359a09de075ae23b4cecb1fb2db97ee2ac6d6

Observation ae84f17b-566d-4557-b30f-08976440d674 · outbound

This paper cites On transforming reinforcement learning with transformers: The development trajectory,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On transforming reinforcement learning with transformers: The development trajectory,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.450134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.723470Z digest=sha256:132b4a36a2dd1f99fb0b6e284ffd03a87bb43bd84cf9addfe7327da8680fd823

Observation a562b3df-f1db-4c05-8ff3-3066d67b3fb7 · outbound

This paper cites Transformer in reinforcement learning for decision-making: A survey,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformer in reinforcement learning for decision-making: A survey,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.437267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.727574Z digest=sha256:3ea26cf596a1ff68da2c103b6a7131e40344f925dee0f0ab2bb636addb102fa9

Observation bc4f5edf-f3a8-418b-9dd3-a8428fb3d894 · outbound

This paper cites Attention Is All You Need.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.731661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.731661Z digest=sha256:7477e11bcd87dfa8e1fd561fa4fc8778f42263f4ba4da459b017a83fefaf1b9c

Observation dc739bf0-97d2-4002-b7ab-733b2d16428b · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.735633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.735633Z digest=sha256:3fdb0f2fd5f1a05bd492c54b2bfc908eecf40b45a0b80dcdfdf6a8956566bf3f

Observation 4cb33299-03b4-4e57-a132-5d852f3352db · outbound

This paper cites Deep Attention Recurrent Q-Network.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Deep Attention Recurrent Q-Network

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:23.170540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.739898Z digest=sha256:d60aa4062855499cb90d3684b1a9e4b158cf6b5a2e39997dbb31e60ca8086e90

Observation bdde2bf2-25da-4f13-825a-1a06159bdf88 · outbound

This paper cites Towards Interpretable Reinforcement Learning Using Attention Augmented Agents.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Towards Interpretable Reinforcement Learning Using Attention Augmented Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.744141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.744141Z digest=sha256:fc5d48e681fccf304136298fea1d7d431530da759cf498087022aa66af56f62c

Observation 15e4264b-f93a-4ea3-9c95-268ffa9b797b · outbound

This paper cites Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.748705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.748705Z digest=sha256:9a7e66292c401814083132de8f8ed3a62f483817e6a3355a7125b63dc3f7bb32

Observation c8c00b89-1879-40dc-97da-39507316d0b6 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.753113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.753113Z digest=sha256:c39445666fb3327a4377ce0cdfa36d6d9d87985d5ce2364a41bd2c80bbdc085e

Observation 95495ef3-5df6-49fb-82f3-4c1d5f2e8526 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Offline reinforcement learning as one big sequence modeling problem,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.422707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.757384Z digest=sha256:925db6b4ac4eba750296d169a04d7f6eb4d5ec09dd4bc83f51cf76061717b01f

Observation cf540214-cd83-474a-abbf-f3ea2833c154 · outbound

This paper cites Online Decision Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Online Decision Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.761745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.761745Z digest=sha256:9be8fb885c15ded2043610e1abf68d2f6e1c4197f630c194d5edbfe81deb0ddd

Observation 1d2e5ca5-fe82-49a4-ab95-acd8ef48894f · outbound

This paper cites Structured State Space Models for In-Context Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Structured State Space Models for In-Context Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.766001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.766001Z digest=sha256:8bda1c7a5bd05f84bfc7e20a08f7d67c6168bc7b9c0628d862706a852af6750d

Observation 9266ec07-5a79-46c4-9b8d-d0c3732ee6af · outbound

This paper cites Mastering Memory Tasks with World Models.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering Memory Tasks with World Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.770341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.770341Z digest=sha256:2f012d5e7bddd89f33024f49db87cd2b31c7b5ebe25eb6cabf80093fcfc30602

Observation 53be7564-2bce-4acd-a07d-f7e7840a8e45 · outbound

This paper cites Mastering atari with discrete world models,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering atari with discrete world models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.408184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.774501Z digest=sha256:411f2a7a65b37fe2f2739ba9b8cab87680951edc82e65c0438be4c5c581bed7f

Observation 95412e92-4c5f-4e87-961d-c47b4c45e61f · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.783296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.783296Z digest=sha256:322287ce9305463f8c05d64231e9a5a5da9da21a4b048edcab2b17a61a1403f0

Observation 1fb5fdc0-918f-403d-803b-cb9169442c2c · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes RWKV: Reinventing RNNs for the Transformer Era

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.787610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.787610Z digest=sha256:9a61a710071175a96b765d8ea979fb2fb9c9fae74fe388b0bf7967a86ece1eca

Observation c4b344d7-2a02-4d2c-8cbe-5b654b71387a · outbound

This paper cites Resurrecting Recurrent Neural Networks for Long Sequences.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Resurrecting Recurrent Neural Networks for Long Sequences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.792509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.792509Z digest=sha256:8e584be6192d20125a1d10ea804a228d60dc1943e18db84d54405c0d3d3bcb96

Observation 16c939f0-da6f-4dd8-88c1-87f85629d068 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficiently modeling long sequences with structured state spaces,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.394090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.797991Z digest=sha256:a74590475359f51ff63963ded4509187376842cfb06cc2c661c7139969691c88

Observation fec57d69-189a-4b98-b57d-14c3d8f1049d · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Simplified State Space Layers for Sequence Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.806620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.806620Z digest=sha256:4a593f64ffa3f4f0e10a306fea23d7299478b17dc56233ec536694f165000850

Observation 28bbaeba-5c66-4e6f-9f28-eae008db609f · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.811361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.811361Z digest=sha256:b5e4ef73bc96a02f386318ab35e5be4a029e2fbe596c74333085cf5190c3d0fd

Observation 30d1836c-1d8b-4175-bf0b-aa61e69e536f · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Efficiently Modeling Long Sequences with Structured State Spaces

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.802215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.802215Z digest=sha256:106178cc6b98f4218e7dd076ebdbae6d0a8c7f13a7e4016c09b9f4851c720887

Observation 97ea078e-db2f-4e07-a3f1-83beccc4c463 · outbound

This paper cites Longformer: The Long-Document Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Longformer: The Long-Document Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.820257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.820257Z digest=sha256:db890f172bb6bd32cfa35cfd5f216232e4c9185b981b58cea2a44d28e67944b9

Observation 83ca39f0-2267-4829-b533-4e45e76c6fca · outbound

This paper cites Transformer-XL: Attentive language models beyond a fixed-length context,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Transformer-XL: Attentive language models beyond a fixed-length context,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.380533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.824663Z digest=sha256:337063f17badc9e081ee612487415bbed77767431dd3a69b324eb80eba18ee3f

Observation 1e28d407-380d-484e-9aca-ba9d95119dfe · outbound

This paper cites Stabilizing Transformers for Reinforcement Learning.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Stabilizing Transformers for Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.815814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.815814Z digest=sha256:992173c3e7ff1adae12cf90822547dff43000ab4277ae9d6186209065f14227f

Observation bfcfe6a5-0e3c-4837-889d-77b8519d1689 · outbound

This paper cites Recurrent Memory Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Recurrent Memory Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.833253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.833253Z digest=sha256:c875cf8727162dc8aceecb80f6bf2da84d396922879b7562de7928adb34480d4

Observation abedf9b8-7a34-40b0-b995-1881e58d5461 · outbound

This paper cites Reformer: The Efficient Transformer.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Reformer: The Efficient Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.837715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.837715Z digest=sha256:902b9d54fd128856132e5a096ace721e19430a48f26c3245656a579cde71f25a

Observation bdddb050-6020-4a30-b036-b7c46fca5a9f · outbound

This paper cites Linear transformers are secretly fast weight programmers,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Linear transformers are secretly fast weight programmers,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.366622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.829200Z digest=sha256:35a3f4ff1cc01b57e2fd628fdca3b2ef40c9020ce14200aa78552ceedae9e322

Observation c186d155-232d-4e77-ac6e-696a6e491d05 · outbound

This paper cites Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Investigating the Role of Feed-Forward Networks in Transformers Using Parallel Attention and Feed-Forward Net Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.846703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.846703Z digest=sha256:c431cdbe8a2aa04a6ae93d6043a23e523ecbabafadc0519fcc58854a347d0d71

Observation 2057806e-73f7-49fb-943d-4dc4d76b0a5e · outbound

This paper cites Rethinking transformers in solving pomdps,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Rethinking transformers in solving pomdps,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.351037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.851223Z digest=sha256:efae309cf595857bdd210fe726381b603931d10bf83dfd6ac1dd6456d78ce124

Observation a7a99566-99ec-4cd8-8365-799994a5d2cb · outbound

This paper cites Rethinking Attention with Performers.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Rethinking Attention with Performers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.842169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.842169Z digest=sha256:19817919e1e4d76e465c407143e4213f636d90bcf32ced56217b3d580ef45c86

Observation 0d42377e-36e6-426d-a376-46d3fb013090 · outbound

This paper cites Pomdp robot domains,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Pomdp robot domains,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.322069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.859622Z digest=sha256:6455668bf586c711a2b2aed624d09075d61ef1d148201c5259d4b8ebed187667

Observation 85361ea8-fcb7-41ea-82e3-c8634bc66ce6 · outbound

This paper cites Learning policies for partially observable environments: Scaling up,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Learning policies for partially observable environments: Scaling up,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.307086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.864054Z digest=sha256:9e1396051f2ade8c730396ee6928b2cae6b676292abb2776b58e0ee4625570da

Observation 1f01bb95-6aeb-41f4-8076-af9d547f5e1c · outbound

This paper cites gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.337172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.855486Z digest=sha256:0807623e43c19e1cdb427e62ec5b624e0dbc7624aed4570379df16e2a4834dd2

Observation 083fdfd6-0e13-46c0-a7d8-14ce4a391f09 · outbound

This paper cites Solving large pomdps using real time dynamic programming,.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Solving large pomdps using real time dynamic programming,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:23.292669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:01:22.868400Z digest=sha256:6868eb510aa5d1ed9bf363f93e3b1c92f6eb50fd8eadbe043ceac83bcb4bc9d1

Observation 1933d1b8-cf1c-4040-9f44-594c111672b1 · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes On Improving Deep Reinforcement Learning for POMDPs

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.710231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.710231Z digest=sha256:c317d31e436824de5b031d0633eae95e390bffdd3889494d5ded3090a8adb99c

Observation 926b6d1d-17ff-4c89-b562-24cb62e6ba2a · outbound

This paper cites Mastering Atari with Discrete World Models.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Mastering Atari with Discrete World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.778724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.778724Z digest=sha256:875dd6614309cb958e80fb77150655f77cf07eed07c1012ccfc8af4835d0c02e

Pith citing papers

No inbound Pith citation observations are available.