AI red teaming plan builder for LLM agents
Describe your AI system and get a scoped AI red teaming plan: the test cases that apply, ranked by risk, each mapped to the OWASP Top 10 for LLM Applications 2025 and MITRE ATLAS, with a coverage score as you run them.
Free, no sign-up. Runs in your browser; nothing is sent to your system or to us.
Your AI red team plan
OWASP Top 10 for LLM Applications 2025 coverage
| Category | Tests | Run | Failed | Note |
|---|
Test cases
Record a result for each test as you run it; the coverage score above updates. Run tests only against systems you own or are authorised to test.
How the plan is built
The catalogue holds 37 test cases across all 10 categories of the OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project), using 34 MITRE ATLAS (Adversarial Threat Landscape for AI Systems) techniques. Your answers decide which apply: a plain chat assistant can't be tricked into deleting files, so it doesn't get that test, while an agent with write access does.
- Priority = the test's base severity (1 to 3) + data sensitivity (0 to 2) + exposure (0 to 2). P1 is 5 or more, P2 is 3 to 4, P3 is 2 or less.
- Coverage = the share of the plan's priority weight you have run, with P1 tests counting three times and P2 twice as much as P3.
- Each test case has steps, an example probe you can paste, and a pass criterion, so two testers judge the same result the same way.
Red teaming the whole system, not one prompt
Testing a system prompt against prompt injection is one test. A red team plan also covers what the system's tools can be made to do, what its retrieval can leak, how its output is rendered, which third-party models and servers it trusts, and what an attacker can cost you. Read what AI red teaming is, how to red team AI agents, or compare open-source AI red teaming tools to run the tests.
Questions
What does the AI red teaming plan builder do?
You describe your AI system (its type, what it can do, the data it can reach and who can use it) and it selects the test cases that apply from a catalogue of 37, ranks them P1 to P3, maps each to the OWASP Top 10 for LLM Applications 2025 and to MITRE ATLAS, and tracks your coverage as you record pass or fail for each test.
Is it free? Do I need an account?
It's free, with no sign-up, and the whole plan is shown on the page. You can copy any test case, or download the plan as Markdown or CSV.
Does it send anything to my AI system or to you?
No. The planner runs entirely in your browser and makes no network requests; the site's content security policy blocks it from sending what you enter anywhere. You run the tests yourself, against systems you own or are authorised to test.
How is this different from a prompt injection test?
A prompt injection test checks one system prompt against injection attacks. This builds a red-team plan for the whole system across all ten OWASP LLM risk categories: tools, retrieval, memory, output handling, supply chain, cost and more.
How is coverage scored?
Each test's priority comes from its base severity plus the sensitivity of your data and how exposed the system is. Coverage is the share of the plan's priority weight you have run (P1 counts three times a P3); tests you mark not applicable are removed from the plan.
Are the OWASP and MITRE ATLAS IDs accurate?
OWASP IDs are from the 2025 edition of the Top 10 for LLM Applications; ATLAS technique IDs and names were checked against ATLAS release 2026.09 on 1 October 2026. Where no ATLAS technique fits a test closely, we leave it unmapped.