Skip to content
Media 365
Book a call
Technology & Platforms

AI Governance, Evals & Assurance

Most AI projects do not fail on capability — they fail on control. We build the evaluation, guardrail and monitoring layer that lets an AI system stay in production safely.

5
core deliverables
5
technologies & channels
1
proof point
What's included

Automated evaluations

Test suites that score accuracy, tone and refusal behaviour on every change, so regressions are caught before customers see them.

Guardrails & policy

Input and output filtering, scope limits, escalation rules and refusal behaviour agreed with your risk owners.

Human-in-the-loop design

Deciding which decisions an agent may make alone, which require review, and how exceptions are routed.

Observability & audit

Tracing every prompt, retrieval and tool call, with logs your auditors and privacy officer can actually read.

Kill switches & rollback

The ability to disable or revert an agent in seconds, tested before launch rather than during an incident.

How we approach it

Industry data through 2026 is blunt about the risk: Gartner expects more than 40% of agentic AI projects to be cancelled by 2027, driven by unclear ROI and weak risk controls, and reported rollback rates fall from roughly 47% for agents without automated evaluations to around 9% with full eval coverage. We treat evaluation and governance as part of the build, not a later phase — and we align delivery to ISO/IEC 42001-style AI management practices and Australia's Voluntary AI Safety Standard.

Technologies & channels
Automated evalsGuardrailsTracing & observabilityHuman-in-the-loopAudit logging

Capability is no longer the constraint

In 2026 almost nobody fails at AI because the model was not clever enough. They fail because nobody could say what good looked like, nobody noticed when it degraded, nobody had authority to switch it off, and nobody could explain to an auditor what it had done. Gartner's forecast that over 40% of agentic AI projects will be cancelled by the end of 2027 attributes the cause to escalating cost, unclear business value and inadequate risk controls — not to technical limits.

Industry reporting through 2026 also puts the governance gap at roughly 60% of organisations running agentic AI without adequate controls in place. That is the gap this service exists to close, and it is far cheaper to build during a project than to retrofit after an incident.

Evaluations: the strongest predictor of survival

An evaluation suite is a fixed set of real cases with known good answers, scored automatically on every change. It sounds mundane and it is the difference between an AI feature that improves and one that quietly rots. Reported rollback rates fall from around 47% for agents with no automated evaluation coverage to roughly 9% with full coverage — a gap large enough to decide whether a project survives.

We build evals from your material, not generic benchmarks: the awkward questions, the edge cases, the queries where the correct response is a refusal. We score accuracy, tone, adherence to policy and refusal behaviour, and we run them in the deployment pipeline so a regression blocks a release rather than reaching customers. When a model provider updates, you upgrade on evidence rather than hope.

Guardrails, and deciding who is actually in charge

A prompt is guidance, not a control. Real guardrails are implemented: input and output filtering, hard scope limits, permission boundaries on every tool an agent can call, escalation paths, and refusal behaviour agreed with the people who own the risk. We run a scoping session with your risk, legal and operational stakeholders, write the policy down in plain language, then implement it so compliance is demonstrable rather than asserted.

Human-in-the-loop design is part of the same conversation, and it is more nuanced than on or off. Different decision types warrant different thresholds — a draft email needs less oversight than a refund, an eligibility determination or anything affecting a ratepayer's account. We set those thresholds per decision, document them, and make the intervention path easy enough that staff actually use it.

Observability, audit and the ability to stop

Every prompt, retrieval and tool call is traced, so an incident can be reconstructed rather than debated. Logs are written to be read by your privacy officer and auditors as well as engineers, which is a deliberate design constraint — governance evidence nobody can interpret is not evidence.

And every system we deploy has a tested kill switch and rollback path. Not a documented intention: an actual mechanism, exercised before launch, that disables or reverts an agent in seconds. The first time you need it should never be the first time it is used.

This scales down as well as up. A single customer-facing assistant still needs a defined scope, refusal behaviour, logging and an off switch — proportionate control, not enterprise ceremony imposed on a small deployment.

Questions we get asked

Why does an AI project need evals?

Because a model update or prompt change can silently degrade quality. Automated evaluations score every change against a fixed set of real cases, which is the single biggest predictor of whether an agent survives in production — reported rollback rates drop from about 47% without eval coverage to roughly 9% with it.

Who decides what the AI is allowed to do?

You do. We facilitate a scoping session with your risk, legal and operational owners, write the policy down, then implement it as enforceable guardrails rather than instructions in a prompt.

Does this apply to a small deployment?

Proportionately, yes. A single customer-facing assistant still needs a defined scope, refusal behaviour, logging and a way to turn it off. We scale the control layer to the risk.

Related work
KYSaaS & Platforms
Know Your Asset (KYA)
When the product promise is “real time”, latency is the product.

Let's talk about AI Governance, Evals & Assurance.

Book a consultation →