Pith. sign in

Paper Citation Record · LEDGER

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents

As of 11 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2501.05501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05501 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:20:01.500799Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b0f7a53-54a8-4bcd-b3be-61cd51e7b03c · outbound

This paper cites Never Give Up: Learning Directed Exploration Strategies.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Never Give Up: Learning Directed Exploration Strategies

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.400140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.400140Z digest=sha256:9a87eb3c2ef07f62e2071923e43be486b5aa53a042bf07416e356ebf8fb87762

Observation 5b3bc92a-e2ae-4d1f-ae4b-7f916ae0e398 · outbound

This paper cites OpenAI Gym.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents OpenAI Gym

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.406481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.406481Z digest=sha256:47e1e95b46b1f90fc27274b379740b1b6c7d20c0a32bfe2ac9d6bc7ccb79a28a

Observation 96f02a0f-43f7-4ea9-9b63-4ec2c605b1c2 · outbound

This paper cites Brown and T.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Brown and T

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.412153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.412153Z digest=sha256:b00d4eb25d259264720d4d03dd9e9136ec0a629dc3ad120172fa9082e5590a1b

Observation 9c5ad7fb-b1d6-48a7-8f75-2530a001e4cb · outbound

This paper cites Brown and T.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Brown and T

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.418213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.418213Z digest=sha256:58fd5de4d5c839fd46ef6d0467787523e1e550d2405797709a86a49dfa653dbc

Observation 4387aa5f-df0a-4520-a506-5d6d572db801 · outbound

This paper cites Dietterich.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Dietterich

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:01.825165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.423972Z digest=sha256:fa5364e54ffcc5a6a61b5870c276475254645b54745dba6dde0c8cf2d52eefee

Observation 353d1950-63f9-4b0b-bacd-0b98ae2168e8 · outbound

This paper cites Deep Recurrent Q-Learning for Partially Observable MDPs.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Deep Recurrent Q-Learning for Partially Observable MDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.429083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.429083Z digest=sha256:a36e8ca2fd688bebfa410e6cd1137812fa8a7d57d524d6cfa761e0aa9a59763f

Observation 7134c656-60c3-436d-9273-2643819e0329 · outbound

This paper cites Hochreiter and J.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Hochreiter and J

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.434743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.434743Z digest=sha256:47eb2ede8697cf4cc67043a60397a6139ec2fb096d32d3f34c3be9249e76e904

Observation 22847d35-5279-4f1a-92bd-16e6a63b1977 · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:20:01.809236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.439472Z digest=sha256:caf8c6ebc16ac8efe34e6e889694044fa3f29111626285865de1c0d71b323e5a

Observation 09188085-a1a5-41f1-83ea-699c165066e5 · outbound

This paper cites Jaakkola, M.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Jaakkola, M

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:01.792040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.444126Z digest=sha256:bb44d069d41f23bfe0a94c213ddc575e5b65ab548049f82121e35fbf224b5c30

Observation af2a37f1-e073-4d29-9beb-e8beaba00405 · outbound

This paper cites Juozapaitis, A.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Juozapaitis, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:01.772961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.448926Z digest=sha256:0a28e8c2cca2753eed3e8327677da443eb464728c08e377a3ef4aa041c87f257

Observation 59d64928-dbf3-4040-9f88-f8c5abfd838f · outbound

This paper cites Karlsson.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Karlsson

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:01.754630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.454330Z digest=sha256:8b109d41a7c65fb90f5c788f2a11b091810ccec4eff8382875a73305a0c2d047

Observation ed44d714-c9fe-4cc4-a7b9-e42c80817f31 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Playing Atari with Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.459523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.459523Z digest=sha256:0bbbc7660074b6332c14b3bacdce5bae362bc249844319a2d58930c402aa639a

Observation 4ca63208-5a39-437a-8245-0b4eef556b3f · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:20:01.738730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.465238Z digest=sha256:42e547dbf226386fed0f2d7d531dbb7b7f4fed88486d340f1ca9ca38bbc77e6f

Observation 8416c1d6-f018-4c69-8693-b0bf36f10b97 · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 14

Resolution
verified exact
doi, observed 2026-08-10T21:20:01.550257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.470221Z digest=sha256:aea1f4de9cec3c18fd6f702949aea689c6f0d6516606423886d6127ee639366a

Observation cb072f77-52b8-4b67-bafb-2735f4b3cea1 · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:20:01.723608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.475461Z digest=sha256:b1d0f1a73d40d091c6af04a28c9c14e90b6c7db95d8e093e7dcb8402f33af2f8

Observation cd345702-7763-42ca-b332-6ecaaa48dbed · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:20:01.707576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.480560Z digest=sha256:cabcac5cbcccc29bdc53a4af1290f8cd1cade74b8b32805f85dbd7f2281cb123

Observation f324169c-03ed-4096-8f9f-fd7393fb20ce · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:20:01.690780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:20:01.485378Z digest=sha256:ee21efe17df4b09ba36bbd170df25e16d3d9f16bcc5c3fa16090d39f328e233c

Observation 55632a50-b590-42cd-b58f-ec5c5c28b98a · outbound

This paper cites Hybrid Reward Architecture for Reinforcement Learning.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Hybrid Reward Architecture for Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.490092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.490092Z digest=sha256:4aa50aace7405c04daa19e3f75268f951be3862b63f444e6d280b48ae988e460

Observation b84584a4-f9ca-4dae-b5f8-b56e23fd1925 · outbound

This paper cites Vinyals, I.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Vinyals, I

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.495703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.495703Z digest=sha256:40bdde5f57422e01209ecdf7eb9a27e78fc5277890bd8420ecefc103306e3938

Observation d720c3d6-e188-43ad-8953-50f1a06cea0e · outbound

This paper cites an unresolved cited work.

Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:01.500799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:20:01.500799Z digest=sha256:892ffca723fcd24fbe44dda642bc68a5ae44efca371b9bbeaeef45f42c57b8de

Pith citing papers

No inbound Pith citation observations are available.