Skip to main content
This is a research prototype. The data and analyses are preliminary and not yet validated — we'd welcome your .

Adversarial input

Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data

Marchal & Xu (2024)

Sub-category
Risk Domain

Vulnerabilities that can be exploited in AI systems, software development toolchains, and hardware, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

"Adversarial Inputs involve modifying individual input data to cause a model to malfunction. These modifications, which are often imperceptible to humans, exploit how the model makes decisions to produce errors (Wallace et al., 2019) and can be applied to text, but also to images, audio, or video (e.g. changing pixels in an image of a panda in a way that causes a model to label it as a gibbon).6"(p. 8)

Part of Misuse tactics to compromise GenAI systems (Model integrity)

Other risks from Marchal & Xu (2024) (22)