Jailbreaks and Prompt Injections Threate…

BackExploiting Limited Generalization of Safety Finetuning

“Model Psychology” Attacks

Home/Risks/Anwar et al. (2024)/Exploiting Limited Generalization of Safety Finetuning

Jailbreaks and Prompt Injections Threate…

“Model Psychology” Attacks

Home/Risks/Anwar et al. (2024)/Exploiting Limited Generalization of Safety Finetuning

Jailbreaks and Prompt Injections Threate…

“Model Psychology” Attacks

Exploiting Limited Generalization of Safety Finetuning

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Anwar et al. (2024)

Sub-category

Risk Domain

2Privacy & Security

2.2AI system security vulnerabilities and attacks

Vulnerabilities that can be exploited in AI systems, software development toolchains, and hardware, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. using encoded text (Wei et al., 2023c) or low-resource languages (Deng et al., 2023a; Yong et al., 2023) (see also Section 3.2)."(p. 69)

Entity— Who or what caused the harm

Human

Due to a decision or action made by humans

AI system

Due to a decision or action made by an AI system

Other

Due to some other reason or is ambiguous

Intent— Whether the harm was intentional or accidental

Intentional

Due to an expected outcome from pursuing a goal

Unintentional

Due to an unexpected outcome from pursuing a goal

Other

Without clearly specifying the intentionality

Timing— Whether the risk is pre- or post-deployment

Pre-deployment

Occurring before the AI is deployed

Post-deployment

Occurring after the AI model has been trained and deployed

Other

Without a clearly specified time of occurrence

Part of Vulnerability to Poisoning and Backdoors

Other risks from Anwar et al. (2024) (26)

Agentic LLMs Pose Novel Risks

7.2 AI possessing dangerous capabilities

AI systemOtherPost-deployment

Multi-Agent Safety Is Not Assured by Single-Agent Safety

7.6 Multi-agent risks

OtherOtherOther

Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs

4.0 Malicious Actors & Misuse

HumanIntentionalPost-deployment

Corporate power may impeded effective governance

6.1 Power centralization and unfair distribution of benefits

OtherUnintentionalOther

Jailbreaks and Prompt Injections Threaten Security of LLMs

2.2 AI system security vulnerabilities and attacks

OtherOtherOther

Vulnerability to Poisoning and Backdoors

2.2 AI system security vulnerabilities and attacks

HumanIntentionalPre-deployment

View all 26 risks from this paper →