Pith. sign in

Paper Citation Record · LEDGER

A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2006.14171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.14171 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:37:43.440877Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:41.546939Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 652b256a-eaaf-4233-94de-631fc1e5f332 · inbound

Effective Analog ICs Floorplanning with Relational Graph Neural Networks and Reinforcement Learning cites this paper.

Effective Analog ICs Floorplanning with Relational Graph Neural Networks and Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T16:43:54.097758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:43:54.097758Z digest=sha256:cf6c669c125c2a72dad00dc734d957024c70abf7470ecfd4b209769332f8c03f

Observation 810ba853-4e80-4603-a081-6ade4818f69b · inbound

Integrating Transit Signal Priority into Multi-Agent Reinforcement Learning based Traffic Signal Control cites this paper.

Integrating Transit Signal Priority into Multi-Agent Reinforcement Learning based Traffic Signal Control A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:20:45.751735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:20:45.751735Z digest=sha256:75f090d56e5d4c44ce58b005351f56db776b597af65a84b2f66da364f0b6e388

Observation c50f1d80-fa3f-41ac-af74-e2ad62556fc9 · inbound

Action Mapping for Reinforcement Learning in Continuous Environments with Constraints cites this paper.

Action Mapping for Reinforcement Learning in Continuous Environments with Constraints A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T21:39:11.945279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:39:11.945279Z digest=sha256:c86e842d57b9b2b197d21f3838b56df9d1f3a932077eedbd1f9c432ea178f9a3

Observation 3793f95a-11f8-46cf-86c2-36f4bc144295 · inbound

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing cites this paper.

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:48.369414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:48.369414Z digest=sha256:0ac57161fce85d874f41545fdbd1da7edc8a85b76f8d6c95783ab17024b69914

Observation 53be11a9-015f-4c00-9348-8634fb2cef0c · inbound

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems cites this paper.

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:53.981879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:53.981879Z digest=sha256:4e140de222102f41a88bde141dc7533d312732142f4d13532bbe66ca7b2bf1b1

Observation 07e50d5f-b39a-42f3-8389-29f39ca7fc18 · inbound

Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems cites this paper.

Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:00.444992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:00.444992Z digest=sha256:68cfdffcdaa870968ec4e2f9728e9ce9481cb5655ac3060d66c772992947a874

Observation 5f748538-feec-4615-9928-e9c37290ccb1 · inbound

Reinforced Context Order Recovery for Adaptive Reasoning and Planning cites this paper.

Reinforced Context Order Recovery for Adaptive Reasoning and Planning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:23:21.079193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:23:21.079193Z digest=sha256:7d0c0f5a30106abba7eeb6405c4aae8e99c00c4828d8b4f1dd175019f7cbfb4e

Observation cbbb56f0-4485-47d6-b0bd-f13aa40aa6e1 · inbound

A Hierarchical Signal Coordination and Control System Using a Hybrid Model-based and Reinforcement Learning Approach cites this paper.

A Hierarchical Signal Coordination and Control System Using a Hybrid Model-based and Reinforcement Learning Approach A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:37:43.440877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:37:43.440877Z digest=sha256:52982b85c17d6d05752a8b04753fde12006baba91129be87eb78f1ea6b932d83

Observation d0ad9eb3-73f8-444f-8359-c85db362407c · inbound

Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609 cites this paper.

Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609 A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:32:07.511541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:32:07.511541Z digest=sha256:c5b1fa3c7f6bd7e2807098e5d615a01a5e325381743753feb245dc6f33ec6868

Observation fe0e8f53-bd8b-43a7-9692-8648a2b0ea01 · inbound

Towards Scalable O-RAN Resource Management: Graph-Augmented Proximal Policy Optimization cites this paper.

Towards Scalable O-RAN Resource Management: Graph-Augmented Proximal Policy Optimization A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:46:46.718720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:46:46.718720Z digest=sha256:503e451f1fa9ee57247849d123e46f7aee7e50fa0351ede0f3127fc5bbe191df

Observation b0b57996-e29b-451e-8ec2-88e49c4e411e · inbound

TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management cites this paper.

TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:14.767914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T01:25:48.075081Z digest=sha256:9a2e00dc5d2237ef3ed12cd9e1a1ea086f00815a7dc51562d7185ee0d6b43ecf

Observation 0a71177f-d0ee-4528-9be5-39d1d5221667 · inbound

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools cites this paper.

Your Loss is My Gain: Low Stake Attacks on Liquid Staking Pools A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:16:11.282290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T17:57:37.260492Z digest=sha256:8d8807ef96b1b9f62d22180abd38f12d58b8a6f68029e3de3789854321b3c370

Observation 4f5e053d-38c2-49ea-9cfe-528aecb8588a · inbound

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency cites this paper.

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.808054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:50:20.040271Z digest=sha256:91e244c7f8e75b37bab8cb48b5683adc1a4eca8dcc7ffa2ce585d166fc3dcbba

Observation 6a89b502-5468-4592-97a4-a75c77dbc6b8 · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:27:41.299409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T17:26:04.762451Z digest=sha256:4dfdb34e3ec761573adfc9ffbcf83cbf7229ed8f9d93ea96ff58930b0495f198

Observation 7f968414-1195-4db5-977e-e03a6be706bb · inbound

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning cites this paper.

Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinforcement Learning A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:15:04.608625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T21:08:47.084947Z digest=sha256:b73398cc7de6e8164a9022da00b310ff488c5e0f143602f4cd9bf082c2ac583c

Observation 67e32243-3306-42fd-971d-72c4e23371c6 · inbound

AlphaTransit: Learning to Design City-scale Transit Routes cites this paper.

AlphaTransit: Learning to Design City-scale Transit Routes A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.317535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T11:55:20.454865Z digest=sha256:1861e80a79aec9c0da8ecf6dd06db6954e0e06d2574ea87f14467e0396a46ace

Observation 7293490f-2098-479d-9857-e34000b2ab40 · inbound

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets cites this paper.

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.596865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:36:53.866075Z digest=sha256:c0d805020809aadbd4317862c49a04158ed0ee58ed21b5ff2f0fd32d337795c8

Observation a3538820-553e-4d4c-ae3c-5905c9db9c56 · inbound

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission cites this paper.

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.548227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T11:25:57.471803Z digest=sha256:ce3b979621b807f7be30209e0de5c8a93177186a7fae697753ded7f4cc1a6d90

Observation 934b253b-b71e-4296-9722-fe10141e3034 · inbound

Optimal Reward Shaping: Autonomous Car Parking Case Study cites this paper.

Optimal Reward Shaping: Autonomous Car Parking Case Study A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T17:40:25.172202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T17:40:25.172202Z digest=sha256:c8711c3a177d77d5c29bdebc057a3750d48a5d26b48e1b0b39e0a0f685eb25ad

Observation 244ead7a-3a31-461e-aa01-b6ae85d09cb4 · inbound

AlphaG-OPD: Reliability-Gated Sibling Counterfactuals for On-Policy Distillation in Symbolic Alpha Factor Discovery cites this paper.

AlphaG-OPD: Reliability-Gated Sibling Counterfactuals for On-Policy Distillation in Symbolic Alpha Factor Discovery A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:13:06.990618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:13:06.990618Z digest=sha256:a0dfc34a3466b4a3be4b89b28afcebdaa632a7871d7d2b3a76e689f71ea8587d