AI Red Teaming Toolsby Agent Trust Cloud

AI red teaming tools

The open-source tools teams use to attack their own LLM applications, compared on facts taken from each project's own repository. Checked 1 October 2026.

The tools at a glance

ToolMaintainerLanguageLicenceLatest releaseBest for
garakNVIDIAPythonApache-2.0v0.17.0 (9 September 2026)Command-line scans of a model or chat endpoint
PyRITMicrosoftPythonMITv1.1.0 (4 September 2026)Scripted, multi-turn attack campaigns written in Python
promptfoopromptfooTypeScriptMIT0.123.1 (18 September 2026)Declarative test suites that run in CI on every change
GiskardGiskardPythonApache-2.0giskard-checks/v1.0.4 (14 September 2026)Evaluation and test checks inside a Python project
DeepTeamConfident AIPythonApache-2.0v1.0.9 (12 November 2025)Python red-team runs against defined vulnerability types
Adversarial Robustness Toolbox (ART)Trusted-AI (LF AI & Data)PythonMIT1.20.1 (7 July 2025)Attacks and defences for models you train

Licence, language and release are read from each GitHub repository. Projects move fast; check the repository before you adopt one.

What each one does

garak

Describes itself as an LLM vulnerability scanner: it runs probes against a model or endpoint and reports which ones got through.

Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM09:2025 Misinformation.

PyRIT

The Python Risk Identification Tool for generative AI: a framework for security professionals to build and automate attacks, including multi-turn ones, and score the results. The project moved from the Azure organisation, whose copy is now archived, to microsoft/PyRIT.

Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM07:2025 System Prompt Leakage.

promptfoo

Tests prompts, agents and RAG pipelines, with red-teaming and vulnerability scanning; configuration is declarative and it runs from the command line and in CI/CD.

Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM06:2025 Excessive Agency, LLM07:2025 System Prompt Leakage.

Giskard

An open-source evaluation and testing library for LLM agents (the repository is now named giskard-oss).

Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM09:2025 Misinformation.

DeepTeam

A framework to red team LLMs and AI agents: it generates attacks against named vulnerability types and evaluates the responses. The repository was still receiving commits on our check date; its latest tagged GitHub release is older.

Most relevant OWASP categories: LLM01:2025 Prompt Injection, LLM02:2025 Sensitive Information Disclosure, LLM06:2025 Excessive Agency.

Adversarial Robustness Toolbox (ART)

A library for machine-learning security covering evasion, poisoning, extraction and inference attacks on classic ML models. Useful when the system includes models you train yourself, less so for prompt-level testing of an LLM app.

Most relevant OWASP categories: LLM04:2025 Data and Model Poisoning, LLM10:2025 Unbounded Consumption.

How to choose

Where our plan builder fits

The AI red teaming plan builder on this site is ours (Agent Trust Cloud). It is not a scanner and sends nothing to your system. It turns a description of your system into a scoped list of 37 possible test cases, each mapped to the OWASP Top 10 for LLM Applications 2025 and to MITRE ATLAS, so you know what to run with the tools above and can track coverage as you go.

Build a red-team plan for your own system in two minutes: Open the free plan builder

Questions

What are the best open-source AI red teaming tools?

Widely used open-source options include garak, PyRIT, promptfoo, Giskard, DeepTeam. They differ in language, how you run them (command line, Python library, CI) and what they focus on, so the right one depends on your stack and on which risks are in scope.

Do AI red teaming tools replace manual testing?

No. Scanners find known attack patterns quickly and are good for regression tests in CI. Business-logic problems, such as an agent that can be talked into acting for the wrong user, usually need a person who understands the system.

Is the plan builder on this site a scanner?

No. It does not send anything to your AI system. It builds a scoped test plan, with test cases mapped to OWASP and MITRE ATLAS, which you can run by hand or with the tools on this page, and it tracks your coverage as you record results.

Sources