AI Red Teaming Toolsby Agent Trust Cloud

Red teaming AI agents

Once a model can call tools, the worst case stops being a bad answer and becomes a bad action. Red teaming an agent means testing what it can be made to do, and by whom.

Five ways agents fail that chatbots don't

  1. Instructions arrive from content, not the user. The agent reads web pages, emails, files and tool responses. Any of them can carry instructions (AML.T0051.001 LLM Prompt Injection: Indirect; LLM01:2025 Prompt Injection).
  2. Tools have more power than the task needs. An agent that only has to read a calendar is often given a token that can also write and delete (LLM06:2025 Excessive Agency; AML.T0053 AI Agent Tool Invocation).
  3. High-impact actions run without a person approving them. Sending, paying, deleting and sharing should show the exact action and wait.
  4. Tools and servers change underneath you. A third-party tool or MCP server can alter its description or behaviour after you approve it (AML.T0109 AI Supply Chain Rug Pull, AML.T0110 AI Agent Tool Poisoning; LLM03:2025 Supply Chain).
  5. Memory and loops make one mistake last. A poisoned memory (AML.T0080.000 AI Agent Context Poisoning: Memory) repeats in every later session; a loop with no budget runs up cost (AML.T0034.002 Cost Harvesting: Agentic Resource Consumption; LLM10:2025 Unbounded Consumption).

A method for agents

  1. Draw the agent's reach

    List every tool, the credential behind it and what that credential allows. List every outside source the agent reads. This map is the attack surface.

  2. Test the joins first

    Plant instructions in each outside source and each tool response, then give the agent an ordinary task. Check which tools it calls.

  3. Test identity

    As a low-privilege user, try to get the agent, or another agent it delegates to, to act with more authority than you have.

  4. Test the brakes

    Confirmation steps, allow-lists, step and spend budgets, and read-only configuration. Try to get around each one.

  5. Log and replay

    Keep each attack as a test case and rerun the set whenever the model, prompt or tool list changes.

The agent test cases in our catalogue

16 of the 37 test cases in the plan builder only apply to systems with tools, memory, code execution or delegation:

Build a red-team plan for your own system in two minutes: Open the free plan builder

Questions

Why is red teaming an AI agent different from testing a chatbot?

A chatbot can only say things; an agent can do things. Once a model can call tools, the worst outcome moves from a bad answer to a bad action: an email sent, a record deleted, data posted to an outside address. Red teaming an agent therefore focuses on tools, permissions, confirmation steps and content the agent reads from outside.

What is indirect prompt injection?

Instructions hidden in content the agent processes, such as a web page, an email, a document or a tool response, rather than typed by the user. MITRE ATLAS lists it as AML.T0051.001 and OWASP covers it under LLM01:2025 Prompt Injection.

How many tests does an agent red team need?

It depends on what the agent can do. The plan builder on this site selects from 37 test cases; a tool-using agent with write access, outside content and personal data typically gets most of them, while a plain chat assistant gets far fewer.

Sources