Find agentic risks before a real adversary does.
Successive waves of autonomous attackers, each wave learning from the last, grounded in the agent's own tools, policies, and personas. Not a library of canned prompts against the model underneath.
The agent you approved is not the agent running
Model updated, tools extended, prompt edited by someone shipping a fix, none of it re-approved. Meanwhile the attack surface is not the model but the whole system around it — its tools, its memory, its multi-turn reasoning, and the policies governing a workflow of several agents. A static library tests what somebody already thought of, which is the part an adversary will not bother with.
“A clean report from a scan everyone runs tells me the agent survived the attacks that are easy to find.”
“I need to know whether it holds up when someone tries to break it, not whether it was configured correctly.”
“A separate team runs it on their schedule and hands me a PDF. By then I have shipped twice.”
Even for the attacks a library does cover, testing everything costs more than anyone will pay, so coverage gets traded against budget. What survives that trade is security theater.
Ten ways a test of an agent can be weaker or stronger
Eight of these are the dimensions from our systematization of the field, published whole. Apply them to whatever you run today, including to this. Most tools are strong on some and weak on others, and a tool that claims all ten is telling you about its marketing rather than its method.
Attackers that learn between waves, not a battery that repeats
A set of attacker agents runs successive waves, each one reading the results of the last. Because the attacks are grounded in the agent — its policies, its tools, its skills — they are specific rather than generic, and they reach the indirect channels a multi-agent system exposes. Coverage compounds, or it repeats; this is the difference.
Coverage is tracked against a taxonomy you choose: the Vijil trust taxonomy of three dimensions, nine categories and thirty-two leaf properties, a published standard such as OWASP, or your own catalog. Adversarial intensity escalates on your budget and your risk tolerance, not ours.
Shaped by the agent and its context.
New attacks generated in the run, from the run.
Evidence a stakeholder can read, and a run id that re-runs it.
In your VPC or fully on-premise, agent agnostic, one container.
One wave, five moves, then it learns
Given a target endpoint, an agent profile, and a risk taxonomy, each wave generates fresh attack seeds, dispatches a multi-turn attacker per seed, judges every transcript against the agent's actual policies, and reflects on what came back. Then it does it again, better.
Attacks and judgments are shaped by the agent's own tools, policies, and personas.
Fresh strategies from what the run has already learned — exploiting what worked, exploring what has not been tried.
Full conversations against the live agent, across direct and indirect surfaces.
Rolls back a stalled turn and sharpens the message when a line of attack stops working.
Reflects on each wave to sharpen the next, until coverage saturates and no new findings emerge.
Waves run until coverage saturates and nothing new comes back. Every finding carries forward to the controls that follow.
60.9% against 39.8%, and here is everything wrong with that number
Share of 266 DecodingTrust-Agent tasks on which Adaptive Red Teaming landed a successful attack within 50 victim messages, against Promptfoo run with its Crescendo multi-turn strategy. Replaying the benchmark's stored attacks reached 44.0%.
One seed. That makes this a suggestive result and not a measured gap, and we would rather you heard it from us. The arms also differ in information access, internal models and compute, so this compares complete configurations rather than isolating attack strategy on its own.
Target: gemini-3.1-pro-preview on the DecodingTrust-Agent harness, undefended
Primary comparison uses the 266 tasks all three arms scored
Between knowing the agent exists and trusting it in production
The agent is found and given an identity it can prove.
Adaptive Red Teaming runs here. Can this agent survive a real adversary?
Findings become the controls Dome enforces at runtime.
Darwin mutates the agent against what the waves found.
Verify answers the question a deployment decision actually turns on. Not whether the agent was configured correctly — whether it holds up when somebody tries to break it. The report and the traces are the evidence the next step is built on.
Point it at an agent and read what comes back
It runs unattended against real load, with rate limiting, retries and a model configured per role. Ship it as one container, through the agent ADK or a long-running A2A server, in your VPC or fully air-gapped. It does not care whose agent framework you chose.
See how your agent holds up against an adaptive adversary
One agent, one run, one id you can hand to whoever asks.