REVIEW 6 cited by
Unrestricted Adversarial Examples
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce a two-player contest for evaluating the safety and robustness of machine learning systems, with a large prize pool. Unlike most prior work in ML robustness, which studies norm-constrained adversaries, we shift our focus to unconstrained adversaries. Defenders submit machine learning models, and try to achieve high accuracy and coverage on non-adversarial data while making no confident mistakes on adversarial inputs. Attackers try to subvert defenses by finding arbitrary unambiguous inputs where the model assigns an incorrect label with high confidence. We propose a simple unambiguous dataset ("bird-or- bicycle") to use as part of this contest. We hope this contest will help to more comprehensively evaluate the worst-case adversarial risk of machine learning models.
Forward citations
Cited by 6 Pith papers
-
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
A conditional generator operating in neural audio codec latent space produces targeted adversarial audio examples in one forward pass, reaching up to 99% success rate at sub-7 ms inference.
-
Position: Adversarial ML for LLMs Is Not Making Any Progress
The authors argue that LLM-era adversarial machine learning is less well-defined, harder to solve, and harder to evaluate, so meaningful progress may not be achievable or trackable in the current paradigm.
-
Generalizability vs. Counterfactual Explainability Trade-Off
A geometric probability called epsilon-VCP rises as models overfit, and the paper offers it as a label-free overfitting diagnostic.
-
Developing Creative AI to Generate Sculptural Objects
ADD and PDD generate 3D point-cloud sculptures by applying DeepDream-style gradient updates to trained classifiers and amalgamating or partitioning point clouds to avoid sparsity.
-
VENOM: Text-driven Unrestricted Adversarial Example Generation with Diffusion Models
A text-to-image diffusion attack that interleaves denoising with momentum-based adversarial gradients and an adaptive on/off switch to generate natural-looking unrestricted adversarial examples.
-
Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters Augmentation
A black-box face recognition attack that augments surrogate models with diverse parameter initializations and hard-model feature perturbations, achieving higher transferability.
Discussion (0). Continue with ORCID to comment.