Pith. sign in

Paper Citation Record · LEDGER

Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2110.05038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.05038 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:48.448037Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:49:02.344960Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 69a54379-6bff-43dd-99dd-511513960743 · inbound

Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids cites this paper.

Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:48.448037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:48.448037Z digest=sha256:da848ceb44d604fe6fbaea64bbd70033c9e4d5b58403bece7c9442c87f5394b8

Observation 1bba64b6-349b-4b40-a804-ee262217b0ec · inbound

GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation cites this paper.

GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T20:48:53.701392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:48:53.701392Z digest=sha256:e4118938f8d571bca8a4fbaf334f0d5513a71e511d3e3958d609cf2e1876f657

Observation ac32d414-1f4a-4eb3-ab82-9f50314332d1 · inbound

Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives cites this paper.

Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T22:51:56.016926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:51:56.016926Z digest=sha256:9c8d8e3010af1aca6806c66cd767ed016e508ba7b58d3276d79897fb121a1553

Observation 5225a48a-9bbf-4634-ae6d-deddcbac7be3 · inbound

MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active Elicitation cites this paper.

MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active Elicitation Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:17:30.678350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:13:42.705619Z digest=sha256:dbb62ffd8d3896b91fe0c7b8e0ac729673968df0885d002676ae14169314e36b

Observation 957fdfbf-93a6-49c2-9e88-592239812121 · inbound

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent cites this paper.

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:46:35.740096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:46:15.275441Z digest=sha256:e90515dd8984c2c05fee526d79073d1b48f1708381d253b5470d6eaec9668823

Observation 24db73d1-e528-4a8c-b857-aa1dcdd88a04 · inbound

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games cites this paper.

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:30:17.383704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:30:17.383704Z digest=sha256:ece50ee75c5440b8564ee78efa9d068687f91d0d726c34bd95627f3d330a6263

Observation ef71cc07-f3e7-44d3-8589-4e31c3b17b07 · inbound

Belief-State RWKV for Reinforcement Learning under Partial Observability cites this paper.

Belief-State RWKV for Reinforcement Learning under Partial Observability Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.937670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:56:09.642064Z digest=sha256:d9233bcd4c917a22f8b6f558bebb00a4bcda87a516f4aebf87a7e9a5e60fc0b8

Observation 5ff5e595-5504-49ed-a64d-a8d53a613bdc · inbound

Recurrent Deep Reinforcement Learning for Chemotherapy Control under Partial Observability cites this paper.

Recurrent Deep Reinforcement Learning for Chemotherapy Control under Partial Observability Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:38.523134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:51:16.753971Z digest=sha256:660a5124ebdd8adc9d988691f2266fd6921c7d2b3839b5c970198ebcdd4c229d

Observation 9a3c8b48-9bf3-44c3-8636-152afeb3f5c2 · inbound

Learning POMDP World Models from Observations with Language-Model Priors cites this paper.

Learning POMDP World Models from Observations with Language-Model Priors Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:57:52.987389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:57:28.642081Z digest=sha256:f4e1855123ef12ccc30c54d81c92a0dc520a9cf8dd0d7eb2224618e892b3997b

Observation 1076a5da-7386-47bf-a13f-c3efd2ab27bb · inbound

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets cites this paper.

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:02.347183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:47:16.490817Z digest=sha256:0fe8a27920bd084c275306b69a6d4ed0c34e316bb46ff14bcae99e43705b28f9

Observation aade19e8-2102-4b5c-b4e0-7afadf6e6cfe · inbound

When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary cites this paper.

When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:40.649329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:40.649329Z digest=sha256:a633b8026b7b6e0496fbe4357551010870b54ac977e178451b6c86f3f6838439