Technical vulnerabilities (Robustness - …

BackTechnical vulnerabilities (Robustness - vulnerability to jailbreaking

Technical vulnerabilities (The risk of m…

Home/Risks/G'sell (2024)/Technical vulnerabilities (Robustness - vulnerability to jailbreaking

Technical vulnerabilities (Robustness - …

Technical vulnerabilities (The risk of m…

Home/Risks/G'sell (2024)/Technical vulnerabilities (Robustness - vulnerability to jailbreaking

Technical vulnerabilities (Robustness - …

Technical vulnerabilities (The risk of m…

Technical vulnerabilities (Robustness - vulnerability to jailbreaking

Regulating under Uncertainty: Governance Options for Generative AI

G'sell (2024)

Source DOI

Sub-category

Risk Domain

2Privacy & Security

2.2AI system security vulnerabilities and attacks

Vulnerabilities that can be exploited in AI systems, software development toolchains, and hardware, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

"Individuals can manipulate models into performing actions that violate the model’s usage restrictions—a phenomenon known as “jailbreaking.” These manipulations may result in causing the model to perform tasks that the developers have explicitly prohibited (see section 3.2.1.). For instance, users may ask the model to provide information on how to conduct illegal activities— asking for detailed instructions on how to build a bomb or create highly toxic drugs."(p. 62)

Entity— Who or what caused the harm

Human

Due to a decision or action made by humans

AI system

Due to a decision or action made by an AI system

Other

Due to some other reason or is ambiguous

Intent— Whether the harm was intentional or accidental

Intentional

Due to an expected outcome from pursuing a goal

Unintentional

Due to an unexpected outcome from pursuing a goal

Other

Without clearly specifying the intentionality

Timing— Whether the risk is pre- or post-deployment

Pre-deployment

Occurring before the AI is deployed

Post-deployment

Occurring after the AI model has been trained and deployed

Other

Without a clearly specified time of occurrence

Supporting Evidence (1)

"Common forms of malicious attacks231 include: • inputting carefully crafted prompts that are able to navigate around a model’s safeguards,232 • extracting training data (especially sensitive information), • backdooring (negating normal authentication procedures to gain unauthorized access to a system), • data poisoning (intentionally compromising a training dataset to manipulate the operation of a model (see below section 3.1.2.B.3.)), and • exfiltration (the theft or unauthorized removal or movement of data).233"(p. 62)

Part of Technical and operational risks