Employ agents that transact in your interest
When software can buy, sell, negotiate and commit on your behalf, reliability is not enough. An agent that works perfectly and answers to the wrong party is worse than one that fails loudly.
A commercial agent is a party to the transaction, not a feature of it
Commerce has always run on delegated authority. We give an agent a mandate, a budget and a boundary, and we hold them to it because they can be identified, because their behavior leaves a record, and because exceeding the mandate has consequences. None of that is true of software by default. An agent that negotiates on your behalf meets counterparties who are optimizing against it, attackers who are probing it, and an incentive to close the deal that nobody wrote down.
A factotum does what it is told. A fiduciary acts in your interest when nobody is checking. The difference is not intelligence — it is accountability, and accountability is built, not promised. The distinction is the one Vijil was founded on: Factotum not fiduciary.
- 01Who is this agent acting for?
- 02What authority has it been delegated?
- 03What information may it disclose?
- 04What transactions may it execute?
- 05How do we know it stays loyal to its principal?
- 06What happens when counterparties and adversaries adapt?
- 07How may it improve without silently changing its mandate?
Seven questions, four answers. Each one is a mechanism you can go and inspect, not a property we claim on the agent’s behalf.
Every agent gets a SPIFFE identity, minted and attested rather than asserted, and an agent card that records what it is for, what it may use, and who answers for it. That is what makes the next three answers enforceable instead of aspirational: you cannot grant authority to a party you cannot name.
Diamond runs adversarial evaluations that negotiate against the agent rather than query it: pressure on price, on disclosure, on the limits of its mandate. Every finding carries the transcript that produced it and the run id that reproduces it, so a failure is a thing you can re-run, not a thing you are asked to accept.
Dome sits in front of the agent as a trusted gateway, uses the attested identity to grant least privilege over tools and data, and decides at runtime what may be disclosed and what may be executed. A transaction outside the mandate does not happen and the attempt is on the record.
An agent scored on outcomes drifts toward whatever the scores pay for, which is how a mandate changes without anyone deciding to change it. Darwin reads the production evidence, proposes a change, and sends it as a pull request that Diamond has already re-verified. The agent improves on your authorization, never around it.
Evidence, not assurance
A trust score across reliability, security and safety, with the findings underneath it and the run id that reproduces each one. A runtime record of what the agent was allowed to do and what it was refused. A pull request for every change, with the evaluation that justified it attached. You should not have to trust us: the taxonomy is published, the detectors are open-weight, and the probes are versioned with their seeds.