What is AI red teaming?
AI red teaming is adversarial testing of an AI system as it is actually deployed: deliberately trying to make it leak data, take actions it shouldn't, or produce output that harms someone, and writing down what worked so it can be fixed.
A working definition
A red team plays the attacker. For an AI system that means sending the model inputs chosen to break it, but also attacking everything around the model: the documents it retrieves, the tools it can call, the memory it keeps, and the screen its answers are rendered on. Most serious failures in AI applications happen at those joins, not inside the model.
The NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024) describes red teaming as one of the ways to test generative AI before and after deployment, alongside evaluations and field testing. Its value is that it looks for the worst case on purpose, where benchmarks report the average case.
How it differs from evaluation and penetration testing
| Benchmarks and evals | AI red teaming | Traditional pen test | |
|---|---|---|---|
| Question | How well does the model do on average? | What is the worst thing someone can make the system do? | Can an attacker get into the infrastructure? |
| Inputs | Fixed test sets | Adversarial, adapted as the tester learns | Exploits against hosts, networks and apps |
| Target | The model | Model plus prompts, retrieval, tools, memory and UI | Servers, APIs, authentication |
| Output | Scores | Reproducible findings with severity and a fix | Vulnerabilities with severity and a fix |
The two security disciplines overlap; see AI penetration testing vs AI red teaming.
The process in six steps
Scope
Write down the system type, what it can do (retrieve, call tools, remember, run code), what data it can reach and who can talk to it. Those facts decide which attacks matter.
Threat model
List who would attack it and why: a curious user, a customer trying to see another customer's data, an outsider planting instructions in a web page the agent reads.
Plan test cases
Pick test cases that match the scope and map them to a shared vocabulary, usually the OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project) and MITRE ATLAS (Adversarial Threat Landscape for AI Systems), so results can be compared and tracked.
Test
Run each case by hand, with tools, or both, in an environment where failures are safe. Keep the exact input and output of every attempt.
Report
Record what failed, how badly, and how to reproduce it. A finding without a reproduction can't be fixed or retested.
Fix and retest
Fix the cause (permissions, filtering, confirmation steps, isolation), then rerun the same cases. Repeat after every significant change to the model, prompt or tools.
Build a red-team plan for your own system in two minutes: Open the free plan builder
What frameworks say
- OWASP publishes the OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project), the ten risk categories most red-team plans for LLM applications are organised around.
- MITRE ATLAS catalogues adversary tactics and techniques against AI systems, with IDs such as AML.T0051.001 LLM Prompt Injection: Indirect (indirect prompt injection), and case studies of real attacks.
- NIST: the NIST AI Risk Management Framework and its NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024) treat adversarial testing as part of measuring and managing generative AI risk.
- EU AI Act: Article 55 of Regulation (EU) 2024/1689 (Artificial Intelligence Act), EUR-Lex requires providers of general-purpose AI models with systemic risk to perform model evaluation that includes conducting and documenting adversarial testing.
Questions
What is AI red teaming?
AI red teaming is structured adversarial testing of an AI system: people (often helped by tools) act like attackers or misusers to find ways the system can be made to leak data, take harmful actions or produce harmful output, so the problems can be fixed before real users or attackers find them.
Is AI red teaming the same as model evaluation?
No. Evaluations and benchmarks measure average behaviour on fixed test sets. Red teaming looks for the worst case: inputs chosen on purpose to break the system as it is deployed, including its tools, data sources and user interface.
Does the law require AI red teaming?
In the EU, providers of general-purpose AI models with systemic risk must conduct and document adversarial testing under Article 55 of the AI Act. For most other systems no law names red teaming, but risk frameworks such as NIST's Generative AI Profile describe it as a way to test and manage risk.
Sources
- OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project)
- MITRE ATLAS (Adversarial Threat Landscape for AI Systems)
- NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024)
- NIST AI Risk Management Framework
- Regulation (EU) 2024/1689 (Artificial Intelligence Act), EUR-Lex
- Checked 1 October 2026.