Pith. sign in

REVIEW 1 cited by

The Dark Side of Explanations: Poisoning Recommender Systems with Counterfactual Examples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.00574 v1 pith:24QHCE3M submitted 2023-04-30 cs.IR

classification cs.IR
keywords counterfactualexplanationsmodelrecommendersystemsh-carsmethodsurrogate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning-based recommender systems have become an integral part of several online platforms. However, their black-box nature emphasizes the need for explainable artificial intelligence (XAI) approaches to provide human-understandable reasons why a specific item gets recommended to a given user. One such method is counterfactual explanation (CF). While CFs can be highly beneficial for users and system designers, malicious actors may also exploit these explanations to undermine the system's security. In this work, we propose H-CARS, a novel strategy to poison recommender systems via CFs. Specifically, we first train a logical-reasoning-based surrogate model on training data derived from counterfactual explanations. By reversing the learning process of the recommendation model, we thus develop a proficient greedy algorithm to generate fabricated user profiles and their associated interaction records for the aforementioned surrogate model. Our experiments, which employ a well-known CF generation method and are conducted on two distinct datasets, show that H-CARS yields significant and successful attack performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vertical Federated Unlearning via Backdoor Certification

    cs.LG 2024-12 reject novelty 3.0 of 10

    A gradient-ascent unlearning algorithm verified by backdoor accuracy is proposed for federated models, but the experimental setup and the unenforced constraint weaken the central VFL claim.

Pith tools