Vijil Diamond
An oversight agent for the agents you are about to trust.

You earned your autonomy. Agents must bring evidence to earn theirs.

An agent is defined by its purpose, the personas of its clients, and the policies of its principal. Diamond compiles the three into a bespoke test suite, runs it against the live agent, and hands back a score with the transcript of what failed.

WHY OVERSIGHT HAS TO BE DELEGATED

Nobody can read every trace

Agents are getting more capable, more autonomous and more numerous at once, and reviewing their behavior by hand stopped scaling. Delegating that judgment moves the need for judgment onto the thing doing the judging: an automated overseer is itself an agent, and it can certify safety it never established. Meanwhile the agent you approved is not the agent running — model updated, tools extended, prompt edited by someone shipping a fix, none of it re-approved. The three people who answer for it feel that differently.

THE BUSINESS OWNER

“Two agents, two vendors, two claims of safety, no common scale. I am choosing on an assertion written by whoever is selling.”

THE RISK OWNER

“I am asked to sign off on an agent whose behavior is described in prose. An approval with nothing behind it is still my signature.”

THE AGENT DEVELOPER

“I fixed what the last review found. Whether the agent is better overall or merely different is not a question my test suite answers.”

A grader that scores its own work drifts, and looks fine doing it
Proxy-scored

When the judge and the judged are tuned against the same signal, the score climbs and the agent stays where it was. The number reports both outcomes identically.

Passed, not probed

A fixed battery run twice gives the same answer twice. The second run measured your memory of the first, not your agent.

All three need the same thing: a score on a published scale, with the failing case behind it.

WHAT DIAMOND IS

Anchor the score to something it never optimizes toward

Diamond is an oversight agent. It compiles the harness from what the agent declares it is for — purpose, personas, policies — rather than from a template, points an adaptive multi-turn adversary at the running system, and scores tool calls and retrievals as well as replies. It returns a trust score in a published taxonomy: reliability, security, safety. The measure itself is anchored to a held-out ground truth Diamond is never tuned against, which is the one part of an oversight agent that cannot be self-certified.

YOUR AGENTSYSTEM UNDER TESTLLM · LLM appcustom agent · multi-agent systemTEST SUITEbuilt-in + bespokepoliciesmetricsprobesdetectorsDiamond test engine87TRUST SCOREadaptive · multi-turn · statefulTRUST REPORTYOUR AGENTSYSTEM UNDER TESTLLM · LLM appcustom agent · multi-agent systemTEST SUITEbuilt-in + bespokepoliciesmetricsprobesdetectorsDiamond test engine87TRUST SCORETRUST REPORTadaptive · multi-turn · stateful
METHODOLOGY

How the score is computed

1
Probes
Attacks and tasks from OWASP LLM Top 10, MITRE ATLAS, garak, and internal red-team.
2
Detectors
Graders label responses pass / fail — calibrated against human-labeled ground truth.
3
Metrics
Pass rates roll up into nine categories, each scored 0–100 with confidence intervals.
4
Score
Weighted by your policy into one Trust Score — traceable to every probe.
REPRODUCIBLEversioned probes + seeds·TRANSPARENTevery score drills to its evidence·CALIBRATEDbenchmarked vs human labels·CURRENTtracks new attacks
Methodology in depth — download the datasheet →·Adaptive offensive AI security testing (PDF) →·Adaptive Red Teaming →·Diamond datasheet →
COMPLIANCE

One policy, tested here and enforced at runtime

Government regulationsEU AI Act · GDPR · NYC LL 144Industry standardsNIST AI RMF · ISO 42001 · OWASPOrganization policiescode of conduct · privacy · terms of useAgent-specific policies✓ permitted✗ prohibitedDiamondcompiles them into bespoke testsCOMPLIANCE REPORTGovernment regulationsEU AI Act · GDPR · NYC LL 144Industry standardsNIST AI RMF · ISO 42001 · OWASPOrganization policiescode of conduct · privacy · terms of useAgent-specific policies✓ permitted✗ prohibitedDiamondcompiles them into bespoke testsCOMPLIANCE REPORT
Dome enforces the same compiled policies at runtime — what Diamond certifies, Dome defends.
WHERE IT FITS

Before you deploy, and again after every change

specYOUR INPUTidentifyDISCOVERverifyDIAMONDdeployYOUR CI/CDdefendDOMEevolveDARWINcodeYOUR AGENTdashed steps use your platform of choice
Diamond is the one test engine in the Trusted Agent Lifecycle, run for three purposes: the developer evaluates, the risk owner verifies, the business owner validates. Dome enforces what Diamond certifies; Darwin evolves what Diamond scores.
WHAT IT COSTS

Three ways in

We publish no list price because there is no meter to read from. The client is free, one agent scored on our hosted console is free, and the paid tier is a deployment inside your own network where you score every agent as often as you like. Every tier is on the pricing page.

Client
Free to install
Free, not open source — Vijil Dome is the Apache-2.0 one.
  • pip install vijil-sdk — the CLI and the Python client
  • The full command surface: evaluate, score history, harnesses, reports
  • The docs, the trust taxonomy, and the rubric our research publishes
  • The evaluation engine runs as the service, because an adaptive swarm does not belong on your laptop
Hosted — free
START HERE
One agent, on our console
  • One agent scored against the Vijil baseline harness
  • Reliability, security and safety, with the cases behind every number
  • Every failure opens to the transcript that produced it
  • Enough to compare one agent you already trust against the scale
Air-gapped
Unmetered, in your VPC
Your hardware is the only limit.
  • Deployed in your own VPC, on-premises, or inside an air-gapped network
  • Every agent, scored as often as you like. No per-evaluation meter.
  • Your policy compiled into the harness, and the multi-turn adaptive adversary
  • Your agents’ transcripts never leave your infrastructure
HOW TO GET STARTED

Register, evaluate, read

1 — Register

Point it at an attested agent. If the agent already carries a workload identity from Vijil Discover, Diamond scores it by that identity — no endpoint, no wrapper, no second registration. An endpoint and a name still work for a one-off, and any agent that speaks chat completions qualifies.

2 — Evaluate

Run the baseline, or point it at your own policy and let it generate a harness from that instead.

3 — Read

The score arrives with the cases that produced it. Open the failures before you argue with the number.

diamond — the whole of it
$ vijil evaluate my-agent# the default taxonomy
$ vijil evaluate my-agent --bespoke# a harness from your policy
$ vijil scores history my-agent# did the agent move, or the test?
WHAT AN EVALUATION RETURNS

A score, a verdict, and the id that reproduces both

A Vijil agent evaluation report: Llama 3.1 8B Instant on Groq, marked FAILED with a trust score of 58.33 against a threshold of 70, over the evaluation id, the harness it ran and the time it was generated.

A number, and the threshold it was judged against. 58.33 against 70 is a failure you can act on; a score with no threshold is a number with an opinion attached.

The evaluation id, the harness and the timestamp travel with it, so the run can be fetched and repeated rather than believed.

VERDICT
FAILED — 58.33 against a threshold of 70.00
AGENT
Llama 3.1 8B Instant, served by Groq
HARNESS
Security
EVALUATION
2cca98e8-5e32-45b1-bac2-2b5399237b99
GENERATED
2026-09-17 18:59:37 UTC
THE RULER, SO YOU CAN CHECK OURS TOO

Where an oversight process is weaker, and where it is stronger

The rubric from our research, published whole. Apply it to whatever you run today, including to Diamond — most tools are strong on some dimensions and weak on others.

Dimension
Weaker
Stronger
Engagement
single-turn probe
multi-turn conversation
Adaptation
static, fixed battery
dynamic, adaptive search
Specialization
generic, one-size suite
custom to purpose, persona, policy
Search
linear, single line
tree: branch, roll back, replay
Coverage
narrow, single harm
holistic, full trust taxonomy
Observation
response-only
trace-aware: tool calls, retrievals
Timing
ahead-of-time only
ahead-of-time and runtime
Anchoring
proxy-scored
held-out ground truth

Underneath it sits the taxonomy the coverage row refers to: three dimensions, nine categories, thirty-two leaf properties, each with its failure modes and root causes. A test that cannot say which leaf it covers is not covering it.

The essay this argument comes from: AI Ethics Benchmarks Are Failing — why a benchmark is not an evaluation →

Score one agent you already trust

Not the one you expect to fail — the one you were sure about.