Guide

AI safety explained: what researchers actually worry about

"AI safety" covers a range of problems, from everyday harms happening now to bigger questions about powerful future systems. Here is a calm overview.

1. Mistakes and hallucinations

Models can be confidently wrong. In low-stakes uses that is an annoyance; in medicine, law or engineering it can hurt people. Mitigations: grounding answers in trusted documents, citations, and keeping humans responsible for decisions.

2. Bias and fairness

Systems trained on historical data can repeat or amplify unfair patterns in hiring, lending or policing. Mitigations: auditing, diverse data, testing across groups, and transparency.

3. Misuse

The same tools that write helpful emails can write scams, disinformation or malware. Developers use usage policies, monitoring, safety training and "red teaming" - deliberately attacking their own systems to find weaknesses.

4. Security of AI itself

AI systems have new attack surfaces: prompt injection (hidden instructions in content the AI reads), data poisoning, and theft of models. As assistants gain the ability to act - send emails, run code, spend money - these attacks matter more, so permissions should be limited and sensitive actions confirmed by a human.

5. Alignment and control

Alignment research asks how to make increasingly capable systems reliably do what we intend, even in situations nobody tested. It includes interpretability (understanding what happens inside models), oversight methods, and evaluating dangerous capabilities before release.

6. Concentration of power and jobs

Who controls the most capable systems, and how work and income change, are social questions as much as technical ones. Policy discussions include transparency rules, audits, and support for workers.

What you can do

  • Stay curious and sceptical of both hype and doom.
  • Keep humans in the loop for important decisions.
  • Protect your own accounts and data.
  • Support transparency and independent testing.

Experts disagree about the likelihood and timing of the larger risks; this guide aims to describe the debate, not settle it.

Keep reading