Pith. sign in

REVIEW 1 cited by

Backdoor Attacks on Multiagent Collaborative Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11455 v1 pith:QAWHUBC6 submitted 2022-11-21 cs.MA

classification cs.MA
keywords adversaryagentothersobservationstriggeractionsaffectagents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Backdoor attacks on reinforcement learning implant a backdoor in a victim agent's policy. Once the victim observes the trigger signal, it will switch to the abnormal mode and fail its task. Most of the attacks assume the adversary can arbitrarily modify the victim's observations, which may not be practical. One work proposes to let one adversary agent use its actions to affect its opponent in two-agent competitive games, so that the opponent quickly fails after observing certain trigger actions. However, in multiagent collaborative systems, agents may not always be able to observe others. When and how much the adversary agent can affect others are uncertain, and we want the adversary agent to trigger others for as few times as possible. To solve this problem, we first design a novel training framework to produce auxiliary rewards that measure the extent to which the other agents'observations being affected. Then we use the auxiliary rewards to train a trigger policy which enables the adversary agent to efficiently affect the others' observations. Given these affected observations, we further train the other agents to perform abnormally. Extensive experiments demonstrate that the proposed method enables the adversary agent to lure the others into the abnormal mode with only a few actions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A malicious pre-trained opponent can, through legitimate in-game actions, embed a trigger-activated backdoor into a victim reinforcement learning agent.

Pith tools