Exploiting Limited Generalization of Saf…

Back“Model Psychology” Attacks

Attacking LLMs via Additional Modalities…

Home/Risks/Anwar et al. (2024)/“Model Psychology” Attacks

Exploiting Limited Generalization of Saf…

Attacking LLMs via Additional Modalities…

Home/Risks/Anwar et al. (2024)/“Model Psychology” Attacks

Exploiting Limited Generalization of Saf…

Attacking LLMs via Additional Modalities…

“Model Psychology” Attacks

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Anwar et al. (2024)

Sub-category

Risk Domain

2Privacy & Security

2.2AI system security vulnerabilities and attacks

Vulnerabilities that can be exploited in AI systems, software development toolchains, and hardware, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.

"LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andreas, 2022), or employing various “social engineering” tricks crafted by humans (Wei et al., 2023c) or other LLMs (Perez et al., 2022b; Casper et al., 2023c)."(p. 69)

Entity— Who or what caused the harm

Human

Due to a decision or action made by humans

AI system

Due to a decision or action made by an AI system

Other

Due to some other reason or is ambiguous

Intent— Whether the harm was intentional or accidental

Intentional

Due to an expected outcome from pursuing a goal

Unintentional

Due to an unexpected outcome from pursuing a goal

Other

Without clearly specifying the intentionality

Timing— Whether the risk is pre- or post-deployment

Pre-deployment

Occurring before the AI is deployed

Post-deployment

Occurring after the AI model has been trained and deployed

Other

Without a clearly specified time of occurrence

Part of Vulnerability to Poisoning and Backdoors

Other risks from Anwar et al. (2024) (26)

Agentic LLMs Pose Novel Risks

7.2 AI possessing dangerous capabilities

AI systemOtherPost-deployment

Multi-Agent Safety Is Not Assured by Single-Agent Safety

7.6 Multi-agent risks

OtherOtherOther

Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs

4.0 Malicious Actors & Misuse

HumanIntentionalPost-deployment

Corporate power may impeded effective governance

6.1 Power centralization and unfair distribution of benefits

OtherUnintentionalOther

Jailbreaks and Prompt Injections Threaten Security of LLMs

2.2 AI system security vulnerabilities and attacks

OtherOtherOther

Vulnerability to Poisoning and Backdoors

2.2 AI system security vulnerabilities and attacks

HumanIntentionalPre-deployment

View all 26 risks from this paper →