Blog/Research notes
RESEARCH NOTES

Look Inside Diamond Adaptive Red Teaming

How DART plans multi-turn attacks from an agent’s own profile, runs them in waves against the live agent, and reflects on each wave to sharpen the next.

Vele Tosevski · October 6, 2026 · 4 min read

As AI agents gain greater access to sensitive data and greater autonomy with tools, developer teams must understand their security posture before deployment. Red teaming evaluates this posture by emulating real-world attackers. As agents grow more complex, their evaluation must become at least as sophisticated.

Modern agent architectures combine context tools, hierarchical memory, MCP connections, and delegation chains. Evaluating systems of this complexity requires multi-turn attacks that exploit the agent and its environment by combining several attack techniques into kill chains.

Manual red teaming by human experts is slow, expensive, and becomes quickly outmoded.

Diamond Adaptive Red Teaming for Agents (DART) automates and accelerates the process of finding failures in your agent’s defenses. DART is a multi-agent system that uses multi-turn attacks to test tool-using agents, with seven key design goals:

  1. Attack through the environment. Interacts with RAG, memory, MCP tools, skills, and connected agents where state changes and actions occur.
  2. Bypass defenses with kill chains. Combines direct and indirect multi-channel prompts to induce policy or safety violations across state changes.
  3. Ground attacks in agent profiles. Uses the agent’s specific purpose, system prompt, and capabilities as attack vectors.
  4. Anchor on standard taxonomies. Maps findings to defined frameworks like the OWASP Top 10 for Agentic Applications or custom taxonomies.
  5. Adapt in depth. Dynamically changes conversation tactics mid-attack and reflects on historical wave results to refine strategies.
  6. Adapt in breadth. Balances exploration of unmapped risk areas with exploitation of discovered weaknesses across test waves.
  7. Ensure verifiable findings. Links every finding directly to reproducible conversation logs for inspection and remediation via Darwin.

Architecture

DART evaluates target agents externally via their endpoint, operating either as SaaS or fully on-premises. DART runs a loop of three steps on four inputs:

  • an agent profile,
  • a taxonomy of failure modes,
  • the personas who use the agent,
  • the policies it must follow.

DART operates an iterative loop across waves: (1) Seed Generation generates test instructions across the taxonomy; (2) Attack Execution executes multi-turn conversations; and (3) Reflection analyzes successes/failures to update the agent profile and refine subsequent waves.

Architecture diagram of Diamond Adaptive Red Teaming for Agents (DART). Four inputs on the left (agent profile, taxonomy, personas, policies) feed a loop that repeats once per wave: 1 seed generation, 2 attack execution, which holds multi-turn conversations with the agent under test behind its endpoint (A2A stateful or chat completions stateless), and 3 reflection, which updates the agent profile. The output on the right is a report of vulnerabilities, policy violations, leaked artifacts and successful strategies, each finding linked to its conversations, in PDF, Markdown or JSON.
Figure 1: How DART works. The numbers mark the three steps of each wave.

The agent under test

DART supports two execution modes for target endpoints:

  • Stateless mode: DART manages conversation history, enabling branching to retry turns without consuming extra budget.
  • Stateful mode: The target agent maintains session state, appending incoming messages directly.

The inputs

The left of Figure 1 shows DART's four inputs:

  • Agent profile: Captures system prompts, endpoints, tools, and workflows. Reflection dynamically updates this profile after each wave.
  • Taxonomy: Structures risk categories (e.g., OWASP Top 10) into node structures containing risk types, guiding questions, and target outcomes to generate focused seeds.
  • Personas: The users who interact with the agent, shaping how the attackers behave.
  • Policies: Defines behavioral and safety constraints used to flag policy violations in reports.

1. Seed generation

Seed generation translates taxonomy nodes into actionable attack instructions via three stages: Node Selection (allocates seed budgets across categories), Outcome Generation (defines specific harms based on the agent profile), and Seed Writing (creates multi-route attack instructions).

For example, for an ASI01 risk node, DART might generate an outcome where a support agent leaks customer records, translating into seeds that trick the agent into forwarding data to an unauthorized recipient.

2. Attack execution

Attacker agents execute seeds in parallel across two main phases:

  1. Planning: The attacker selects a persona, context, and multi-turn sequence.
  2. Interaction: Executes the plan dynamically, adapting tactics, branching, or switching strategies based on model responses until the objective is reached or budgets expire.

For example, the support agent would refuse a direct request to email customer records outside the company, so the attacker splits that request into three steps that each look routine. First, posing as a customer, it asks how support cases are handled, and the agent reveals that it reads a records system and forwards cases with an email tool. Next, it asks about a specific case, and the agent retrieves that customer's record into its context. Finally, it pastes in an email that appears to come from a colleague, asking the agent to forward the case to a new address, and the agent emails the customer's record out of the company with its own tool (goals 1 and 2).

3. Reflection

At the end of each wave, reflection extracts successful/failed strategies, updates the agent profile with leaked information (e.g., internal tool names), and logs harms and policy violations to inform subsequent test waves.

The report

Upon run completion, DART produces a report covering vulnerabilities, policy violations, leaked artifacts, and successful strategies. Each finding includes direct transcript evidence and is downloadable in PDF, Markdown, or JSON formats.

Each vulnerability is filed under its taxonomy node with a quote or summary of the strongest evidence, and every entry links to the conversations behind it, whichever wave they ran in (goal 7). The report is available as PDF, Markdown, or JSON.

Getting started: running a red team evaluation

To run an evaluation, navigate to Verify > Create Evaluation in the Console UI, select a target agent, pick an Effort preset, attach personas/policies, and launch. Watch the demo video for a full walkthrough.

Beyond the Console UI

You can also launch evaluations programmatically via Vijil CLI, Python SDK, REST API, or MCP server integration in CI/CD pipelines. Get started by installing the Vijil plugin to Claude Code.

References

← All posts