AI glossary

42 terms, explained without jargon.

Agent
An AI system that can plan and take actions - such as browsing, running code or calling tools - to complete a goal, rather than only answering a single question.
AGI (artificial general intelligence)
A hypothetical AI that can learn and perform most intellectual tasks as well as a human. There is no agreed test or date, and experts disagree on how close current systems are.
Alignment
Making an AI system pursue the goals and values its designers and users actually intend, including avoiding harmful or deceptive behaviour.
API
Application programming interface: a way for software to talk to an AI model over the internet, usually paid per amount of text processed.
Attention
The mechanism in transformer models that lets each part of the input weigh how relevant every other part is. It is why modern models handle context so well.
Benchmark
A standard test used to compare models, such as exams or coding challenges. Scores can be gamed or leaked into training data, so treat them as hints, not proof.
Bias
Systematic unfairness in model outputs, usually inherited from imperfect training data or design choices.
Chain of thought
Having a model write out intermediate reasoning steps before the final answer, which often improves results on maths and logic.
Chatbot
A program that converses in natural language. Modern chatbots are usually built on large language models.
Context window
The amount of text (measured in tokens) a model can consider at once - your prompt, earlier conversation and its own reply together.
Deepfake
Synthetic image, audio or video that convincingly imitates a real person, made with generative AI.
Diffusion model
A kind of generative model that creates images (or audio/video) by gradually removing noise from random static, guided by a text prompt.
Embedding
A list of numbers that represents the meaning of text, an image or other data, so similar things end up close together. Used for search and recommendations.
Fine-tuning
Training an existing model a bit more on specialised data so it performs better at a particular task or style.
Foundation model
A large model trained on broad data that can be adapted to many tasks.
Generative AI
AI that creates new content - text, images, audio, video or code - rather than only classifying or predicting.
GPU
Graphics processing unit: the type of chip whose parallel processing makes training and running large models practical.
Guardrails
Rules, filters and training that keep an AI system within safe and acceptable behaviour.
Hallucination
When a model states something false or invented with confidence, such as a fake citation. Always verify important facts.
Inference
Running a trained model to produce an output (as opposed to training it). Every time you send a prompt you are doing inference.
Jailbreak
A prompt designed to trick a model into ignoring its safety rules.
Large language model (LLM)
A neural network trained on huge amounts of text to predict the next token, which gives it the ability to write, summarise, translate and reason in language.
LoRA
Low-rank adaptation: a cheap way to fine-tune a model by training small add-on weights instead of the whole network.
Machine learning
Teaching computers to find patterns in data and improve at a task without being explicitly programmed with rules.
Multimodal
Able to work with more than one type of data, such as text, images, audio and video.
Neural network
A computing system made of layers of simple connected units whose connection strengths (weights) are learned from data.
Open-weight model
A model whose trained weights are published so anyone can run or adapt them. Not always the same as fully open source, because training data and code may be withheld.
Overfitting
When a model memorises its training examples and performs poorly on new data.
Parameter
A learned number inside a model. Larger models have billions or trillions of parameters, but bigger is not always better.
Prompt
The text (and sometimes images) you give a model to tell it what to do.
Prompt injection
An attack where hidden instructions in a web page, document or email trick an AI assistant into doing something its user did not ask for.
RAG (retrieval-augmented generation)
Letting a model look up relevant documents first and use them to answer, which reduces made-up facts and lets it use fresh or private information.
Reinforcement learning
Training by trial and reward: the system tries actions and learns which lead to better outcomes.
RLHF
Reinforcement learning from human feedback: people rank model answers and the model is trained to prefer the better ones. A key step in making chatbots helpful.
Red teaming
Deliberately attacking or stress-testing an AI system to find safety and security problems before real attackers do.
Synthetic data
Artificially generated data used to train or test models, often produced by other models.
System prompt
Hidden instructions from the developer that set a chatbot's role, tone and limits before your message arrives.
Temperature
A setting controlling randomness: low values give predictable answers, high values give more varied and creative ones.
Token
A chunk of text (roughly three-quarters of an English word on average) that models read and write. Pricing and limits are counted in tokens.
Training data
The examples a model learns from. Its quality, diversity and legality shape what the model can do and how it fails.
Transformer
The neural network design, introduced in 2017, behind nearly all modern language models and many image and audio models.
Vector database
A database built to store embeddings and quickly find the most similar items.