Brainforest × Vijil: trustworthiness in natural language search
About 60% better search accuracy, and end-to-end production deployment in two weeks.
Brainforest developed a natural language search assistant for a leading real estate firm in Sweden on the DigitalOcean Gradient platform. But the reliance on a large language model for search functionality created significant challenges in trustworthiness, particularly regarding the accuracy and reliability of search results.
The problem
Before partnering with Vijil, Brainforest faced a variety of issues that hindered the agent’s performance.
Hallucinations
The system generated extraneous content when essential facts were missing or scattered across multiple data points.
Low accuracy
User searches returned irrelevant listings that were outside of search parameters.
Innumeracy
The inability to perform numerical operations limited the agent’s usefulness for price-based queries and filtering.
Exposure
The agent’s open response to all prompts made it vulnerable to misuse and prompt injection attacks.
These reliability issues would result in users abandoning search queries, reducing customer engagement, while the security vulnerabilities would make the agent susceptible to external attacks.
Vijil’s trust optimization framework
Vijil began by defining “trustworthy” from the user’s perspective, creating a custom yet comprehensive test harness for reliability, security and safety. Vijil then used its test engine to run the harness at scale to produce a trust score and a trust report that assessed the risks and recommended mitigations based on the context and requirements of the agent’s function.
Finally, Vijil implemented the mitigations that Brainforest authorized, enhancing the agent to make it ready for production in a few short weeks. That entire process is now automated and reusable by other customers.
1. Custom test harness development
Vijil created a comprehensive test harness to evaluate the agent’s reliability based on sample queries, security, and safety. This included assessing the risks associated with the model’s outputs, and prompt injection attack testing.
2. Optimizing prompt response accuracy
Recognizing the limitations of the existing knowledge base, Vijil suggested a function-calling approach that transformed user queries into API requests to a live database. This ensured that responses were grounded in real-time data from a current data source.
3. SQL-like query support
The agent was optimized to construct SQL-like queries, allowing it to handle complex numerical queries and respond with accurate data — addressing a key obstacle to reliability.
4. Enhanced security measures
Improvements to the system prompt and guardrails reduced vulnerability to misuse and enhanced the overall security of the agent.
Trustworthiness achieved
The agent could now generate precise answers to complex inquiries, significantly reducing the occurrence of hallucinations. Users could perform numerical operations, such as sorting properties by price. And with controls in place to block adversarial manipulation and enforce system prompt protection, the agent was better protected against prompt injection and vulnerability exploits.
- about 60% improvement in search accuracy;
- reduced hallucinations, giving more reliable results;
- two weeks, end to end, to production deployment.
Read this case study as a PDF →
The same harness runs against your agent. Start for free, ortalk to us about an evaluation against your policy.