RESEARCH

You should not have to trust us

An instrument nobody can inspect is a vendor’s opinion, and an opinion is not evidence in front of a regulator. So the taxonomy is published, the detectors are open-weight and on Hugging Face, the probes are versioned with their seeds, and the method is open for anyone outside this company to check. Credentials ask you to trust us. This page is the argument that you need not. If you answer for what an agent does, that is the difference between a number you can put in front of a regulator and one you cannot.

TOPICS

What we work on

Adaptive red-teaming
Probes and attack strategies that surface hallucination, injection, jailbreak, and leakage failures — before adversaries do.
Trust measurement
Metrics and methodology for scoring reliability, security, and safety — with detectors calibrated against human-labeled ground truth.
Governance to guardrails
Small, fast open models for runtime defense — detecting prompt injections at inference time without adding latency.
Agent adaptation
Methods for attributing production failures to root causes and evolving agents under explicit constraints.
PUBLICATIONS & ARTIFACTS

Published work

PAPERS
What the platform runs on

Each of these underwrites a mechanism you can go and check on a product page.

Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
nonfactuality detection
Lamb, Ivanova, Torr, Rudner (Vijil) · Preprint · arXiv:2604.07172 · 2026
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
nonfactuality detection
Bhardwaj, Kempe, Rudner (Vijil) · ICML 2026 · arXiv:2510.21891 · 2025
Red Teaming AI Red Teaming
the adversarial method
Majumdar (Vijil), Pendleton, Gupta · CAMLIS 2025 · arXiv:2507.05538 · 2025
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
the reliability dimension
Novikova, Anderson, Blili-Hamelin, Majumdar (Vijil) · ICML 2025 workshop (R2-FM) · arXiv:2505.00268 · 2025
Improving Consistency in Large Language Models through Chain of Guidance
the reliability dimension
Raj, Gupta, Rosati, Majumdar (Vijil) · TMLR 2025 · arXiv:2502.15924 · 2025
Embedding-Based Classifiers Can Detect Prompt Injection Attacks
Dome’s runtime detectors
Ayub, Majumdar · CAMLIS 2024 · arXiv:2410.22284 · 2024
Is ETHICS About Ethics? Evaluating the ETHICS Benchmark
why a score is not evidence
Hancox-Li (Vijil), Blili-Hamelin · Preprint · arXiv:2410.13009 · 2024
Evaluating Defences Against Unsafe Feedback in RLHF
the safety dimension
Rosati, Edkins, Raj, Atanasov, Majumdar (Vijil), et al. · AAAI 2025 workshop (AICS) · arXiv:2409.12914 · 2024
garak: A Framework for Security Probing Large Language Models
Diamond’s probe framework
Derczynski, Galinkin, Martin, Majumdar (Vijil), Inie · Preprint · arXiv:2406.11036 · 2024
Representation Noising: A Defence Mechanism Against Harmful Finetuning
the safety dimension
Rosati, Wehner, Williams, et al., incl. Majumdar (Vijil) · NeurIPS 2024 · arXiv:2405.14577
Other work by our people

Published by Vijil researchers, and not part of the platform.

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
Fang, Yuan, Kong, Rudner (Vijil) · ICLR 2026 · arXiv:2607.19575 · 2026
Open Problems in Frontier AI Risk Management
Ziosi, Plueckebaum, Casper, et al., incl. Rudner (Vijil) · Oxford Martin AI Governance Initiative · 2026
Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning
Barazandeh, Majumdar (Vijil), Rajyaguru, Michailidis · ICMLA 2025 · arXiv:2506.00236 · 2025
Stop Treating ‘AGI’ as the North-Star Goal of AI Research
Blili-Hamelin, Graziul, Hancox-Li (Vijil), et al. · ICML 2025 · arXiv:2502.03689
Unsocial Intelligence: An Investigation of the Assumptions of AGI Discourse
Blili-Hamelin, Hancox-Li (Vijil), Smart · AIES 2024 · arXiv:2401.13142 · 2024
Showing all 15 papers · sorted by recency
OPEN MODELS
Prompt-injection detection models
DeBERTa- and ModernBERT-based classifiers
The same detectors Dome runs, released for anyone to use, benchmark or fine-tune. They are the strongest thing on this page: you do not have to take our word about a detector you can download and point at your own traffic.
Download →
BENCHMARK
Dome against five competing guardrails
Llama 3.3 70B Instruct on Groq · Nvidia RTX 4090 · default out-of-box configurations · milliseconds
GuardrailTrust Score upliftp50 callp95 call
GCP Model Armor+19.7227.9543.2
Vijil Dome (Q3 ’26)+17.620.755.9
DeBERTa prompt injection+15.724.147.9
AWS Bedrock Guardrails+14.0320.7495.4
Vijil Dome (Q1 ’26)+13.423.166.2
Nvidia NemoGuard+12.7178.5231.7
Meta PromptGuard 2+3.624.757.1

Where it loses. GCP Model Armor scores a higher uplift than we do, +19.7 against our +17.6, and DeBERTa returns a lower p95 than we do, 47.9ms against our 55.9ms. On this table we are not the top line. On the report’s separate accuracy run over the same 15,000-prompt set, balanced accuracy is 97.7% against GCP’s 72.6% — a different measurement, stated separately because it is measured separately.

Where it wins. The combination. Dome is the most accurate guardrail on the curated set at 97.7% balanced accuracy, 25 points above GCP’s 72.6%. GCP is the one guardrail here that scores a higher trust uplift, and it spends 227.9ms at the median to do it against our 20.7ms. The only guardrail with a lower p95 than ours, DeBERTa at 47.9ms, reaches 62.0% accuracy. A guard in the request path is judged on both, because a fast guard that misses the attack is not protecting anything.

One limitation, stated plainly: the uplift column is measured with our own Trust Score, so we are scoring competitors with our instrument. The latency columns are not ours — they are wall clock, on the same hardware, through the same standard APIs.

More research notes on the Vijil blog.

Try to beat our detector

Benchmark our prompt-injection detector against yours. It is on Hugging Face, it is small enough to run on your own hardware, and we would rather hear where it loses than where it wins — which is the only kind of marketing this page’s readers should accept.