We give your AI a score you can actually read
Accuracy, hallucination, RAG hit rate, regression, cost and risk — a number for each of the six. Start with the free triage, then decide whether to go further.
You might be stuck here
The hard part after AI goes live isn't whether it runs. It's that you can't say how accurate, how stable or how expensive it currently is.
You're going on gut feel
- You can't say whether it's accurate.
- You test one prompt at a time until it feels right.
- Ask someone else and you get a different verdict.
How we handle itThe same test set, run repeatedly, giving you a number you can track.
Afraid to change anything
- You tweak a prompt and don't know what else broke.
- There's no test set you can re-run.
- Every release feels like a fresh bet.
How we handle itYou get a regression script — run it before a release and you know what broke.
No idea what's burning money
- The bill goes up every month.
- You can't see which call costs the most.
- Or which calls don't need to happen at all.
How we handle itCost broken down per call, with the savings pointed out.
The six things we measure
The same test set run repeatedly, each dimension producing a number you can track. Run it again after the next release and you know whether it got better or worse.
Accuracy
- Whether answers are right, and how far off the reference they are.
- Scored per scenario, not one blended average.
- Produces a baseline you can track against.
Hallucination rate
- Whether it invents things that don't exist.
- The rate it happens at.
- With the representative failures attached.
RAG hit rate
- Whether the right passage was retrieved at all.
- Whether the citation points to the right part of it.
- Lists the questions answered wrongly or not at all.
Regression
- Whether this change broke the previous version.
- A re-runnable script delivered to you.
- Run it before every release.
Cost
- Tokens and money per call.
- Identifies the most expensive calls.
- Points out which calls can be merged or dropped.
Risk
- Prompt injection, privilege escalation, data leakage.
- Mapped against OWASP, ISO, NIST and local regulator guidance.
- Delivered as a prioritised fix list.
Which tier to pick
Start with the free triage to see whether going further is worth it. If it is, pick a tier by how deep you need to go.
L1 Standard audit
- You need a baseline number.
- So every future release can be compared against it.
- Includes the regression script and the report.
L2 Deep audit
- You have a baseline and want to know why it's off.
- Includes workflow optimisation diagnosis.
- Includes a light code audit.
L3 Enterprise governance
- You need to map to OWASP, ISO, NIST and regulator guidance.
- Includes a full system and security audit.
- Includes a governance framework and rollout support.
Pricing
The free triage carries no obligation. One-off audits and the monthly health check can be bought separately.
Free triage
Usually includes:One pass over a small set of real questions, giving you the broad picture and a view on whether a full audit is worth it.
Affects the quote:No obligation to continue.
Final price is confirmed once the scope is clear · First consultation is free
L1 Standard audit
Usually includes:Baseline measurement of accuracy, hallucination rate and RAG hit rate, a regression test script, a written report and a walkthrough call.
Affects the quote:Size of the test set, number of scenarios measured.
Final price is confirmed once the scope is clear · First consultation is free
L2 Deep audit
Usually includes:Everything in L1, plus workflow optimisation diagnosis, a light code audit and a prioritised fix list.
Affects the quote:System complexity, whether the code audit is included.
Final price is confirmed once the scope is clear · First consultation is free
L3 Enterprise governance
Usually includes:Full system and security audit, governance framework, compliance mapping (OWASP / ISO / NIST / regulator guidance) and rollout support.
Affects the quote:Compliance scope, number of systems, length of the support period.
Final price is confirmed once the scope is clear · First consultation is free
Monthly health check
Usually includes:Continuous monitoring of hallucination rate, RAG hit rate, regression and tool-call failures, with a monthly report and anomaly alerts.
Affects the quote:Monitoring frequency and number of checks. Three tiers; annual billing is discounted.
Final price is confirmed once the scope is clear · First consultation is free
How it works
Start with the free triage to see whether going further is worth it. If it is, pick a tier by how deep you need to go.
- 01
Free triage
- One pass over a small set of real questions.
- The broad picture and the obvious failure modes.
- A view on whether a full audit is worth it.
- 02
Build the test set
- A test set built from your own real questions.
- Reference answers confirmed with you.
- This same set gets re-run every time from now on.
- 03
Measure
- One pass for each of the six dimensions.
- A number for each.
- Same set, same conditions — so results are comparable.
- 04
Report and fix order
- A written report with a line of interpretation on every number.
- One walkthrough call.
- Plus which items to fix first.
Common questions
What does the free triage actually show?
We run one pass over a small set of real questions and give you an approximate accuracy figure, the obvious failure modes, and a view on whether a full audit is worth it.
Why can't you just guarantee an accuracy figure?
Accuracy is tied to the distribution of your own questions. What we can do is measure where it is now, point out what can be improved, and use the same test set to verify the improvement actually happened.
Will we understand the report?
Every number comes with a line explaining what it means, and there's a walkthrough call.
Do you fix the problems as well?
You get a prioritised fix list. Having us do the fixing is a separate conversation — or your existing vendor can fix it and we re-verify.
Can you audit AI another vendor built for us?
Yes. Start with the free triage for the overall picture, then decide whether a standard or deep audit is worth it.