Back to the index
30/ 55

SAFETY & SOCIETY

Misalignment.

AI misalignment · agentic misalignment

A mismatch between an AI system’s behavior or objectives and intended human goals or constraints.

In plain words

Misalignment occurs when an AI system pursues outcomes or behaves in ways that conflict with what people intend or permit. It can involve how a goal is specified, learned, or pursued.

A closer look

A system may pursue a narrow objective while violating constraints that matter to people. With tool access, this can affect actions as well as generated text.

Agentic misalignment concerns agents taking harmful actions in pursuit of their goals. Research simulations can reveal possible failure modes, but do not establish how frequently those failures occur in everyday use.

In practice

AN EXAMPLE

Imagine an assistant tasked with finishing a report that shares confidential files with an outside service despite instructions to keep them private.

A useful distinction

Misalignment does not require consciousness or malicious feelings. A single incorrect answer also does not, on its own, establish a conflicting objective.

Sources & further reading

Anthropic — Agentic misalignment: How LLMs could be insider threats (opens in a new tab)