Adaptive Red Teaming
Part of Vijil Diamond.

Find agentic risks before a real adversary does.

Successive waves of autonomous attackers, each wave learning from the last, grounded in the agent's own tools, policies, and personas. Not a library of canned prompts against the model underneath.

WHY A LIBRARY OF ATTACKS IS NOT A TEST

The agent you approved is not the agent running

Model updated, tools extended, prompt edited by someone shipping a fix, none of it re-approved. Meanwhile the attack surface is not the model but the whole system around it — its tools, its memory, its multi-turn reasoning, and the policies governing a workflow of several agents. A static library tests what somebody already thought of, which is the part an adversary will not bother with.

THE BUSINESS OWNER

“A clean report from a scan everyone runs tells me the agent survived the attacks that are easy to find.”

THE RISK OWNER

“I need to know whether it holds up when someone tries to break it, not whether it was configured correctly.”

THE AGENT DEVELOPER

“A separate team runs it on their schedule and hands me a PDF. By then I have shipped twice.”

Even for the attacks a library does cover, testing everything costs more than anyone will pay, so coverage gets traded against budget. What survives that trade is security theater.

THE RULER, SO YOU CAN CHECK OURS TOO

Ten ways a test of an agent can be weaker or stronger

Eight of these are the dimensions from our systematization of the field, published whole. Apply them to whatever you run today, including to this. Most tools are strong on some and weak on others, and a tool that claims all ten is telling you about its marketing rather than its method.

Dimension
Weaker
Stronger
System under test
the model
the agent — its tools, memory, and policy
Engagement
single-turn probe
multi-turn conversation
Adaptation
fixed battery
adapts inside the run
Search
linear, one line
tree — branch, roll back, replay
Specialization
generic suite
grounded in the agent's policy and persona
Observation
response only
trace-aware — tool calls, retrievals
Coverage
one harm
the full trust taxonomy
Anchoring
proxy-scored
held-out ground truth
Reproducible
a report
a run id that re-runs
Timing
before deployment
before deployment and at runtime
WHAT ADAPTIVE RED TEAMING IS

Attackers that learn between waves, not a battery that repeats

A set of attacker agents runs successive waves, each one reading the results of the last. Because the attacks are grounded in the agent — its policies, its tools, its skills — they are specific rather than generic, and they reach the indirect channels a multi-agent system exposes. Coverage compounds, or it repeats; this is the difference.

Coverage is tracked against a taxonomy you choose: the Vijil trust taxonomy of three dimensions, nine categories and thirty-two leaf properties, a published standard such as OWASP, or your own catalog. Adversarial intensity escalates on your budget and your risk tolerance, not ours.

Grounded

Shaped by the agent and its context.

Adaptive

New attacks generated in the run, from the run.

Auditable

Evidence a stakeholder can read, and a run id that re-runs it.

Yours to run

In your VPC or fully on-premise, agent agnostic, one container.

AGENTS AGAINST AGENTS

One wave, five moves, then it learns

Given a target endpoint, an agent profile, and a risk taxonomy, each wave generates fresh attack seeds, dispatches a multi-turn attacker per seed, judges every transcript against the agent's actual policies, and reflects on what came back. Then it does it again, better.

1 — Agent profile grounding

Attacks and judgments are shaped by the agent's own tools, policies, and personas.

2 — Attack steering

Fresh strategies from what the run has already learned — exploiting what worked, exploring what has not been tried.

3 — Multi-turn episodes

Full conversations against the live agent, across direct and indirect surfaces.

4 — Adaptive retry

Rolls back a stalled turn and sharpens the message when a line of attack stops working.

5 — Analysis and self-improvement

Reflects on each wave to sharpen the next, until coverage saturates and no new findings emerge.

Waves run until coverage saturates and nothing new comes back. Every finding carries forward to the controls that follow.

MEASURED, ON ONE SEED

60.9% against 39.8%, and here is everything wrong with that number

Share of 266 DecodingTrust-Agent tasks on which Adaptive Red Teaming landed a successful attack within 50 victim messages, against Promptfoo run with its Crescendo multi-turn strategy. Replaying the benchmark's stored attacks reached 44.0%.

One seed. That makes this a suggestive result and not a measured gap, and we would rather you heard it from us. The arms also differ in information access, internal models and compute, so this compares complete configurations rather than isolating attack strategy on its own.

Vijil benchmark DIAMOND-598, September 2026
Target: gemini-3.1-pro-preview on the DecodingTrust-Agent harness, undefended
Primary comparison uses the 266 tasks all three arms scored
WHERE IT FITS

Between knowing the agent exists and trusting it in production

IDENTIFY

The agent is found and given an identity it can prove.

VERIFY

Adaptive Red Teaming runs here. Can this agent survive a real adversary?

DEFEND

Findings become the controls Dome enforces at runtime.

EVOLVE

Darwin mutates the agent against what the waves found.

Verify answers the question a deployment decision actually turns on. Not whether the agent was configured correctly — whether it holds up when somebody tries to break it. The report and the traces are the evidence the next step is built on.

HOW TO GET STARTED

Point it at an agent and read what comes back

diamond
$ vijil register my-agent# once, to give it an identity
$ vijil evaluate my-agent# the waves run here
$ vijil reports my-agent# the findings, and the id that reproduces them

It runs unattended against real load, with rate limiting, retries and a model configured per role. Ship it as one container, through the agent ADK or a long-running A2A server, in your VPC or fully air-gapped. It does not care whose agent framework you chose.

See how your agent holds up against an adaptive adversary

One agent, one run, one id you can hand to whoever asks.