Collusion between LLM-Agents

BackDual-Use Capabilities Enable Malicious Use and Misuse of LLMs

Misinformation and Manipulation

Home/Risks/Anwar et al. (2024)/Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs

Collusion between LLM-Agents

Misinformation and Manipulation

Home/Risks/Anwar et al. (2024)/Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs

Collusion between LLM-Agents

Misinformation and Manipulation

Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Anwar et al. (2024)

Category

"Like all technologies, LLMs have the possibility for misuse by malicious actors. Malicious use of dual- use capabilities of AI is a recurring concern within literature (Brundage et al., 2018; Hendrycks et al., 2023; Mozes et al., 2023)"(p. 84)

Entity— Who or what caused the harm

Human

Due to a decision or action made by humans

AI system

Due to a decision or action made by an AI system

Other

Due to some other reason or is ambiguous

Intent— Whether the harm was intentional or accidental

Intentional

Due to an expected outcome from pursuing a goal

Unintentional

Due to an unexpected outcome from pursuing a goal

Other

Without clearly specifying the intentionality

Timing— Whether the risk is pre- or post-deployment

Pre-deployment

Occurring before the AI is deployed

Post-deployment

Occurring after the AI model has been trained and deployed

Other

Without a clearly specified time of occurrence

Other risks from Anwar et al. (2024) (26)

Agentic LLMs Pose Novel Risks

7.2 AI possessing dangerous capabilities

AI systemOtherPost-deployment

Multi-Agent Safety Is Not Assured by Single-Agent Safety

7.6 Multi-agent risks

OtherOtherOther

Corporate power may impeded effective governance

6.1 Power centralization and unfair distribution of benefits

OtherUnintentionalOther

Jailbreaks and Prompt Injections Threaten Security of LLMs

2.2 AI system security vulnerabilities and attacks

OtherOtherOther

Vulnerability to Poisoning and Backdoors

2.2 AI system security vulnerabilities and attacks

HumanIntentionalPre-deployment

Vulnerability to Poisoning and Backdoors > Natural Language Underspecifies Goals

7.1 AI pursuing its own goals in conflict with human goals or values

OtherUnintentionalPre-deployment

View all 26 risks from this paper →