Back to the index
04/ 55

SAFETY & SOCIETY

AI safety.

Artificial intelligence safety

Research and practices aimed at preventing harmful AI behavior and outcomes.

In plain words

AI safety studies how to build and use AI systems while reducing the risk of harm. It includes improving model behavior, evaluating risks, and controlling how systems are deployed.

A closer look

Safety work addresses both unintended failures and deliberate misuse. Alignment, interpretability, evaluations, and operational safeguards contribute different kinds of evidence and protection.

Safety depends on context: a drafting assistant and an agent with access to important systems can create different risks. Assessments must consider the model, its tools, and how people use it.

In practice

AN EXAMPLE

Before launching an assistant, a team tests harmful requests, restricts sensitive actions, and monitors failures after release.

A useful distinction

AI safety is not a single filter or a guarantee of zero risk. Good performance on ordinary tasks does not prove safe behavior in every setting.

Sources & further reading

Anthropic — Core views on AI safety (opens in a new tab)