Blog/Announcements
LAUNCH ANNOUNCEMENT

Adaptive Red Teaming in Vijil Diamond, Tailored to Your Agent

Adaptive Red Teaming runs multi-turn adversarial tests built from your agent’s own policies, tools, and personas, inside your VPC. On the DTap benchmark it succeeded on 1.38 times as many attack tasks as the benchmark’s own curated attacks.

Vijil · October 8, 2026 · 7 min read

Today we launched Adaptive Red Teaming in Vijil Diamond, an automated testing system that finds security vulnerabilities and policy violations in tool-using AI agents using multi-turn attacks. On a public benchmark for agent security, Adaptive Red Teaming succeeded on 1.38 times as many attack tasks as the benchmark's own curated attacks. You can try Adaptive Red Teaming for free with your agent using a coding agent plugin or through console.vijil.ai. Or sign up to deploy Adaptive Red Teaming as a container alongside your agent in development inside your VPC.

Title card: Vijil Diamond Adaptive Red Teaming, DART. Five numbered turns linked by arrows labeled learns, above the attack surfaces it probes: prompts, tools, memory, identity, data and other agents.

Watch: Adaptive Red Teaming in action

The practice of red teaming helps organizations identify vulnerabilities and weaknesses in their systems by using a structured adversarial approach to probing their defenses. However, a growing number of AI agents are now part of the enterprise environment, and older red teaming tools are falling short of finding the new risks that agents present. At the same time, newer approaches to red teaming AI agents remain complex and expensive because they require expert professional services or specialized tools, both scarce, to keep up with advances in agentic software. The open source tools that scan for vulnerabilities are often research projects that need to be adapted and integrated before an enterprise can use them reliably. This leaves enterprises that need to test their AI agents, within their own corporate network with no prompts or responses leaving the perimeter, searching for a capability that few independent software vendors can offer.

Adaptive Red Teaming helps enterprises test their AI agents built in any framework using any model deployed on any platform of their choice. Adaptive Red Teaming is an automated testing system that runs adversarial evaluations over the whole agent, not only the model inside but also the tools it calls through the MCP servers to which it’s connected and the other agents to which it delegates tasks using the A2A protocol. Because tests are customized to the target agent’s policies, tools, and skills, they are realistic and context-specific. Adaptive Red Teaming finds issues and proposes changes – from hardening the system prompt to modifying control policies – to mitigate the risks. You can run Adaptive Red Teaming in minutes in a container deployed alongside your agent inside your VPC.

The gap between what gets tested and what gets attacked

Most AI red teaming in the market today builds on existing methodologies and testing practices that were useful starting points, but fall short in specific ways when applied to agentic AI.

From vulnerability scanning, AI red teaming inherited the known-signature model (especially when applied to LLM testing): compile a library of attacks, run them, report pass/fail. That works when the target under test is deterministic and known failures are already cataloged. Agents, on the other hand, can refuse a request cleanly on Monday, but may surrender confidential data on Thursday. The failure that matters most is one for which a probe doesn’t yet exist but emerges late in a chained attack. In our benchmark, about a quarter of Adaptive Red Teaming's successful attacks landed after the tenth turn.

From penetration testing, it inherited the discrete-engagement model: a specialist team, a scoped window, a report delivered weeks later. This produces a real artifact, but it sits outside the agent development process and is not easily consumed by AI teams. By the time findings reach the engineers who can act on them, the agent has changed, or worse is already in production. The result is a point-in-time audit of a dynamic system.

From adversarial machine learning, it inherited perturbation techniques that insert noise with stealth into the input so that a change that’s imperceptible to a human would cause the model to fail in its classification task. Even a carefully constructed and calibrated message that launches a successful attack doesn't test what happens when the real-world adversary is patient, adapts, and works across fifty turns instead of one. These single-prompt adversarial tests for AI agents mark the floor, but don’t help agent developers and security teams build realistic defenses. What developers need is the adversarial mindset guiding tests that anticipate the target agent’s controls, adapt to blocked attempts, and persist in probing the agent’s use of tools, the memory it carries between turns, and the other agents to which it hands work.

What Adaptive Red Teaming does differently

Before the first turn, Adaptive Red Teaming plans an attack for each goal: a persona, a context, and a turn-by-turn approach. An attacker model then writes each turn and adapts to the agent's replies. A judge scores progress after every turn, and a reflection step revises the plan when a line of attack stalls.

Cover of a Vijil Red Team Engagement Report, an adversarial safety assessment of the agent sdk-walkthrough-openai-mcp: 23 findings documented across 2 waves of 10 seeds each.

Attacks that worked get pushed further. Attacks that fail are retried with new tactics. You spend the test budget on finding new risks.

We designed Adaptive Red Teaming to deliver greater effectiveness (as measured by coverage and attack success rate) at lower complexity (through automated and customized deployment) at lower cost (compared to consulting services, cybersecurity platforms, and DIY solutions).

Key capabilities:

Tailored to your agent. Before the first attack runs, Adaptive Red Teaming profiles your agent with whatever information is available in the agent card: purpose, system prompt, MCP servers, tool manifest. Adaptive Red Teaming builds its attacks from that profile. Your declared user personas become attacker types. Your tools become the action surface an attack has to work through. Your policies become the specific rule an attack tries to break, and the standard every transcript gets judged against. The same system pointed at a travel-booking agent and a finance-approval agent produces categorically different attacks, because the questions worth asking are different.

Adaptive and multi-turn. Attacks run as full conversations against the live agent, across both direct channels and the indirect ones that matter in multi-agent systems. When a line of attack stalls, the system sharpens and retries rather than moving on. This is how a persistent human red teamer works. Some failures appear only under sustained pressure, and a single-pass scanner does not reach them. You control adversarial intensity, so you can dial testing up or down against cost and your own risk tolerance.

Auditable against a customizable taxonomy. Findings are tracked against a risk category and the outcome that evidences it. Adaptive Red Teaming ships with predefined taxonomies including OWASP's Agentic Security Initiative (ASI) Top 10, and can be pointed at any risk catalog you supply. The output is explicit coverage: what was probed, what held, what didn't, and what hasn't been tested yet.

Built to run where your agents run. Adaptive Red Teaming runs in your environment, so test data stays inside your network. It runs in your VPC or fully on-premise for air-gapped and regulated environments.

How Adaptive Red Teaming compares with a benchmark's own attacks

We tested Adaptive Red Teaming against the stored attacks that come with the DecodingTrust-Agent Platform (DTap), a public benchmark for the security of AI agents. Each DTap task gives an agent a goal, names a harmful action, and includes a stored attack designed to cause it. We ran Adaptive Red Teaming and the stored attacks on the same 275 tasks across 12 business domains. The target was an agent built on Google's Gemini 3.1 Pro with no added defenses. The benchmark's authors optimized each stored attack against a different agent and selected the best by hand. Each stored attack is a single message sent once. Adaptive Red Teaming holds a conversation of up to 50 turns and changes its approach based on the agent's replies.

Adaptive Red Teaming succeeded on 163 of the 275 tasks, or 59.3%. The stored DTap attacks succeeded on 118 tasks, or 42.9%. Adaptive Red Teaming therefore succeeded on 1.38 times as many tasks, with a 95% confidence range of 1.21 to 1.58. Adaptive Red Teaming needed 9 turns to match the stored attacks and built its lead after that point. It led in 8 of the 12 domains and tied in 2. In each of the other 2 domains it lagged by one task. The two methods often succeed on different tasks. We tuned Adaptive Red Teaming on DTap tasks against this same agent before the test. Removing the tasks used in tuning leaves the result between 1.36 and 1.38 times.

Where Adaptive Red Teaming fits in your agent development loop

Adaptive Red Teaming supports a critical stage of the trusted agent loop. An agent enters the testing stage after it is registered, with a persistent identity, an agent card, and governance ownership established. Adaptive Red Teaming reads that agent card; the agent’s purpose, tools, policies, and user personas shape the attack prompts and ground the judge's verdicts. The output of evaluation creates evidence for the next stage, where hardening and runtime controls address the findings.

For an agent developer, Adaptive Red Teaming delivers a structured report with findings, traces, and proposed fixes. Because evaluation is grounded in the agent's own configuration, findings point at specific issues the attack uncovered, along with the transcript of the turn and categorization of the risk. And because it runs unattended against your agent in development, it can run in your CI/CD pipeline.

The Verify screen in the Vijil Console. An agent is chosen from a list of registered agents, and the Adaptive test configuration sets effort (Quick 10 min, Standard 30 min, Deep 60 min), seeds per wave, max concurrency, and optional personas and policies.

For the security engineer or GRC team, the results provide context and evidence a stakeholder can easily grasp. Coverage against a named taxonomy, findings tied to risk categories, and traces that show how each conclusion was reached. Diamond scores agents across reliability, security, and safety dimensions, which gives a deployment decision a measured basis.

A completed adaptive red teaming run: 20 attackers over 107 turns in 29 minutes. Wave 1 lists each seed instruction beside its OWASP ASI risk type, such as tool misuse, insecure inter-agent communication and cascading failures, and the risk outcome it targets.

As AI agents take on greater autonomy within enterprise workflows, static testing and point-in-time audits are no longer enough. Securing your agentic environment requires continuous, adaptive adversarial testing that runs where your data lives.

Stop guessing how your agent will hold up under sustained pressure. Score your agent today to run your first automated red teaming evaluation in minutes, or book a 30-minute walkthrough to see a live demonstration tailored to your agent.

Adaptive Red Teaming is available now. Visit Adaptive Red Teaming.

← All posts