REVIEW 3 cited by
Rethinking Adversarial Attacks in Reinforcement Learning from Policy Distribution Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep Reinforcement Learning (DRL) suffers from uncertainties and inaccuracies in the observation signal in realworld applications. Adversarial attack is an effective method for evaluating the robustness of DRL agents. However, existing attack methods targeting individual sampled actions have limited impacts on the overall policy distribution, particularly in continuous action spaces. To address these limitations, we propose the Distribution-Aware Projected Gradient Descent attack (DAPGD). DAPGD uses distribution similarity as the gradient perturbation input to attack the policy network, which leverages the entire policy distribution rather than relying on individual samples. We utilize the Bhattacharyya distance in DAPGD to measure policy similarity, enabling sensitive detection of subtle but critical differences between probability distributions. Our experiment results demonstrate that DAPGD achieves SOTA results compared to the baselines in three robot navigation tasks, achieving an average 22.03% higher reward drop compared to the best baseline.
Forward citations
Cited by 3 Pith papers
-
Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation
DAC and BP-DAC generate 'unsourced' adversarial CAPTCHAs from semantic prompts and report transfer attack success rates above 95% on ImageNet classifiers in black-box settings.
-
Rethinking Membership Inference Attacks Against Transfer Learning
A white-box attack on the student model can infer teacher-training membership in transfer learning by comparing the student's hidden representations with those of a shadow student model.
-
Secure Resource Allocation via Constrained Deep Reinforcement Learning
A deep Q-network with a fixed deadline penalty is claimed to cut simulated system cost by up to 40% and energy use by 41.5% in serverless multi-cloud offloading.
Discussion (0). Continue with ORCID to comment.