REVIEW 2 cited by
Generating Natural Language Attacks in a Hard Label Black Box Setting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study an important and challenging task of attacking natural language processing models in a hard label black box setting. We propose a decision-based attack strategy that crafts high quality adversarial examples on text classification and entailment tasks. Our proposed attack strategy leverages population-based optimization algorithm to craft plausible and semantically similar adversarial examples by observing only the top label predicted by the target model. At each iteration, the optimization procedure allow word replacements that maximizes the overall semantic similarity between the original and the adversarial text. Further, our approach does not rely on using substitute models or any kind of training data. We demonstrate the efficacy of our proposed approach through extensive experimentation and ablation studies on five state-of-the-art target models across seven benchmark datasets. In comparison to attacks proposed in prior literature, we are able to achieve a higher success rate with lower word perturbation percentage that too in a highly restricted setting.
Forward citations
Cited by 2 Pith papers
-
Confidence Elicitation: A New Attack Vector for Large Language Models
Confidence elicitation, asking a model to verbalize its uncertainty, can be used as a soft-label signal to craft stronger black-box word-substitution attacks on LLMs.
-
TrustGLM: Evaluating the Robustness of GraphLLMs Against Prompt, Text, and Structure Attacks
GraphLLMs are broadly vulnerable to text, graph structure, and prompt label attacks, but the severity depends heavily on the model and dataset.
Discussion (0). Continue with ORCID to comment.