AI red teaming tools
The open-source tools teams use to attack their own LLM applications, compared on facts taken from each project's own repository. Checked 1 October 2026.
The tools at a glance
| Tool | Maintainer | Language | Licence | Latest release | Best for |
|---|---|---|---|---|---|
| garak | NVIDIA | Python | Apache-2.0 | v0.17.0 (9 September 2026) | Command-line scans of a model or chat endpoint |
| PyRIT | Microsoft | Python | MIT | v1.1.0 (4 September 2026) | Scripted, multi-turn attack campaigns written in Python |
| promptfoo | promptfoo | TypeScript | MIT | 0.123.1 (18 September 2026) | Declarative test suites that run in CI on every change |
| Giskard | Giskard | Python | Apache-2.0 | giskard-checks/v1.0.4 (14 September 2026) | Evaluation and test checks inside a Python project |
| DeepTeam | Confident AI | Python | Apache-2.0 | v1.0.9 (12 November 2025) | Python red-team runs against defined vulnerability types |
| Adversarial Robustness Toolbox (ART) | Trusted-AI (LF AI & Data) | Python | MIT | 1.20.1 (7 July 2025) | Attacks and defences for models you train |
Licence, language and release are read from each GitHub repository. Projects move fast; check the repository before you adopt one.
What each one does
garak
Describes itself as an LLM vulnerability scanner: it runs probes against a model or endpoint and reports which ones got through.
Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM09:2025 Misinformation.
PyRIT
The Python Risk Identification Tool for generative AI: a framework for security professionals to build and automate attacks, including multi-turn ones, and score the results. The project moved from the Azure organisation, whose copy is now archived, to microsoft/PyRIT.
Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM07:2025 System Prompt Leakage.
promptfoo
Tests prompts, agents and RAG pipelines, with red-teaming and vulnerability scanning; configuration is declarative and it runs from the command line and in CI/CD.
Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM06:2025 Excessive Agency, LLM07:2025 System Prompt Leakage.
Giskard
An open-source evaluation and testing library for LLM agents (the repository is now named giskard-oss).
Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM09:2025 Misinformation.
DeepTeam
A framework to red team LLMs and AI agents: it generates attacks against named vulnerability types and evaluates the responses. The repository was still receiving commits on our check date; its latest tagged GitHub release is older.
Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM06:2025 Excessive Agency.
Adversarial Robustness Toolbox (ART)
A library for machine-learning security covering evasion, poisoning, extraction and inference attacks on classic ML models. Useful when the system includes models you train yourself, less so for prompt-level testing of an LLM app.
Most relevant OWASP categories: LLM04:2025 Data and Model Poisoning, LLM10:2025 Unbounded Consumption.
How to choose
- You want a quick scan of a model or chat endpoint: a probe-based scanner such as garak.
- You want tests that run on every pull request: a declarative, CI-friendly tool such as promptfoo.
- You need custom, multi-turn attack campaigns: a framework such as PyRIT, written in Python.
- You train or fine-tune your own models: add a classic ML robustness library such as ART for poisoning and extraction tests.
- Whatever you pick: scope first. A tool runs the attacks it knows; it can't tell you that your agent's email tool is the riskiest part of the system.
Where our plan builder fits
The AI red teaming plan builder on this site is ours (Agent Trust Cloud). It is not a scanner and sends nothing to your system. It turns a description of your system into a scoped list of 37 possible test cases, each mapped to the OWASP Top 10 for LLM Applications 2025 and to MITRE ATLAS, so you know what to run with the tools above and can track coverage as you go.
Build a red-team plan for your own system in two minutes: Open the free plan builder
Questions
What are the best open-source AI red teaming tools?
Widely used open-source options include garak, PyRIT, promptfoo, Giskard, DeepTeam. They differ in language, how you run them (command line, Python library, CI) and what they focus on, so the right one depends on your stack and on which risks are in scope.
Do AI red teaming tools replace manual testing?
No. Scanners find known attack patterns quickly and are good for regression tests in CI. Business-logic problems, such as an agent that can be talked into acting for the wrong user, usually need a person who understands the system.
Is the plan builder on this site a scanner?
No. It does not send anything to your AI system. It builds a scoped test plan, with test cases mapped to OWASP and MITRE ATLAS, which you can run by hand or with the tools on this page, and it tracks your coverage as you record results.