You earned your autonomy. Agents must bring evidence to earn theirs.
An agent is defined by its purpose, the personas of its clients, and the policies of its principal. Diamond compiles the three into a bespoke test suite, runs it against the live agent, and hands back a score with the transcript of what failed.
Nobody can read every trace
Agents are getting more capable, more autonomous and more numerous at once, and reviewing their behavior by hand stopped scaling. Delegating that judgment moves the need for judgment onto the thing doing the judging: an automated overseer is itself an agent, and it can certify safety it never established. Meanwhile the agent you approved is not the agent running — model updated, tools extended, prompt edited by someone shipping a fix, none of it re-approved. The three people who answer for it feel that differently.
“Two agents, two vendors, two claims of safety, no common scale. I am choosing on an assertion written by whoever is selling.”
“I am asked to sign off on an agent whose behavior is described in prose. An approval with nothing behind it is still my signature.”
“I fixed what the last review found. Whether the agent is better overall or merely different is not a question my test suite answers.”
When the judge and the judged are tuned against the same signal, the score climbs and the agent stays where it was. The number reports both outcomes identically.
A fixed battery run twice gives the same answer twice. The second run measured your memory of the first, not your agent.
All three need the same thing: a score on a published scale, with the failing case behind it.
Anchor the score to something it never optimizes toward
Diamond is an oversight agent. It compiles the harness from what the agent declares it is for — purpose, personas, policies — rather than from a template, points an adaptive multi-turn adversary at the running system, and scores tool calls and retrievals as well as replies. It returns a trust score in a published taxonomy: reliability, security, safety. The measure itself is anchored to a held-out ground truth Diamond is never tuned against, which is the one part of an oversight agent that cannot be self-certified.
How the score is computed
One policy, tested here and enforced at runtime
Before you deploy, and again after every change
Three ways in
We publish no list price because there is no meter to read from. The client is free, one agent scored on our hosted console is free, and the paid tier is a deployment inside your own network where you score every agent as often as you like. Every tier is on the pricing page.
pip install vijil-sdk— the CLI and the Python client- The full command surface: evaluate, score history, harnesses, reports
- The docs, the trust taxonomy, and the rubric our research publishes
- The evaluation engine runs as the service, because an adaptive swarm does not belong on your laptop
- One agent scored against the Vijil baseline harness
- Reliability, security and safety, with the cases behind every number
- Every failure opens to the transcript that produced it
- Enough to compare one agent you already trust against the scale
- Deployed in your own VPC, on-premises, or inside an air-gapped network
- Every agent, scored as often as you like. No per-evaluation meter.
- Your policy compiled into the harness, and the multi-turn adaptive adversary
- Your agents’ transcripts never leave your infrastructure
Register, evaluate, read
Point it at an attested agent. If the agent already carries a workload identity from Vijil Discover, Diamond scores it by that identity — no endpoint, no wrapper, no second registration. An endpoint and a name still work for a one-off, and any agent that speaks chat completions qualifies.
Run the baseline, or point it at your own policy and let it generate a harness from that instead.
The score arrives with the cases that produced it. Open the failures before you argue with the number.
A score, a verdict, and the id that reproduces both

A number, and the threshold it was judged against. 58.33 against 70 is a failure you can act on; a score with no threshold is a number with an opinion attached.
The evaluation id, the harness and the timestamp travel with it, so the run can be fetched and repeated rather than believed.
- VERDICT
- FAILED — 58.33 against a threshold of 70.00
- AGENT
- Llama 3.1 8B Instant, served by Groq
- HARNESS
- Security
- EVALUATION
- 2cca98e8-5e32-45b1-bac2-2b5399237b99
- GENERATED
- 2026-09-17 18:59:37 UTC
Where an oversight process is weaker, and where it is stronger
The rubric from our research, published whole. Apply it to whatever you run today, including to Diamond — most tools are strong on some dimensions and weak on others.
Underneath it sits the taxonomy the coverage row refers to: three dimensions, nine categories, thirty-two leaf properties, each with its failure modes and root causes. A test that cannot say which leaf it covers is not covering it.
Score one agent you already trust
Not the one you expect to fail — the one you were sure about.