Skip to main content
Start a conversation

Technical guide · Product, engineering, data and risk teams releasing AI-enabled workflows.

Production AI evaluation: evidence before release

Production AI evaluation is the disciplined comparison of a defined system configuration against representative tasks, known failure modes and release criteria. It is not a single benchmark score, a successful demo or a claim that the model is generally reliable.

01

Working position

Evaluate the system that users will encounter

A model is rarely the whole system. Users encounter instructions, prompt templates, retrieval, business rules, tools, interface controls, permissions, human review and fallback behaviour. An evaluation that tests only a base model can be useful research, but it may not answer whether the released workflow works. Identify the versioned configuration: model or service, prompt, retrieval source, tool permissions, policy rules, UI path and environment. That is the thing for which a release decision is being made.

The target should be expressed in user and operational terms. What task is supported, what output is expected, what must never be automated, what evidence must be shown and what happens on uncertainty? This prevents the team from optimising an attractive metric while missing the decision that gives the product its value. It also makes later changes comparable rather than treating every model update as an entirely new unknown.

02

Delivery reasoning

Build a representative, governed evaluation set

Evaluation examples need traceable provenance. Record where each case came from, why it represents the intended workload, whether it includes personal or confidential information, who approved its use and how long it may be retained. Use a mix of ordinary cases, difficult cases, incomplete inputs, conflicting evidence, adversarial or misleading inputs where relevant and situations in which the correct answer is to decline, escalate or ask for more information.

Do not let the evaluation set become a hidden training or prompt-tuning loop without recording it. If examples influence instructions, retrieval or rules, separate development and hold-out cases where practical. Small teams do not need an elaborate laboratory to benefit from this discipline; they do need enough separation to avoid claiming that a system works simply because it was repeatedly tested against cases it was shaped around.

03

Delivery reasoning

Choose measures that reflect the decision and its harms

Accuracy can matter, but it is often incomplete. A classification workflow may need precision and recall by class; a drafting workflow may need reviewer acceptance, traceability and harmful-claim rate; a retrieval workflow may need groundedness, citation usefulness, latency and abstention behaviour. Establish the unit of review and the reviewer rubric. Where expert judgement is required, record how reviewers are calibrated and how disagreement is handled rather than presenting one person's view as an objective benchmark.

Set release criteria before reading the results. They can be thresholds, required controls, unacceptable failure types or a qualitative decision record. The crucial point is that criteria need to be connected to the use case. A strong average can conceal an unacceptable failure in a high-impact route. NIST's AI RMF and playbook are voluntary frameworks, but their emphasis on governing, mapping, measuring and managing risks provides a useful structure for this decision.

04

Delivery reasoning

Test failure paths and operational boundaries

A credible evaluation makes failure visible. Test missing context, ambiguous requests, conflicting instructions, invalid tool inputs, retrieval with no useful result, content that should be refused, outages, slow responses and user attempts to bypass intended controls. If the product uses tools or downstream systems, verify permission boundaries and the result of a partial failure. The correct system behaviour may be a safe stop, a request for clarification or an escalation—not a fluent answer.

For workloads processing personal data or supporting decisions about people, involve the relevant privacy, legal, security or sector specialists early. Evaluation evidence cannot substitute for the underlying obligations. It can, however, make the system's capabilities and limitations more inspectable, which is necessary for a responsible release conversation.

05

Delivery reasoning

Turn results into a release record and a monitoring plan

A useful release record names the configuration evaluated, cases and time period, measures, results, known limitations, approval decision and owner. It should state whether the result supports a limited release, further iteration or a stop. Attach evidence in a format the reviewers can inspect without exposing sensitive evaluation data unnecessarily. A short decision table is often more valuable than a large dashboard with no explanation of the data behind it.

Release is not the final measurement. Define what will be observed in use, how user feedback and exceptions are collected, what drift or regression signal matters, who reviews it and which actions are possible. If a model, prompt, retrieval source or tool behaviour changes, decide whether the change is minor, requires targeted regression checks or needs a new release review. The answer is workload-specific; the record makes that judgement repeatable.

06

Decision record

Make each decision inspectable before the work moves on.

01

What configuration is under evaluation?

Versioned model, prompt, retrieval, tools, UI and policy controls. Accountable owner: Technical/product owner. Boundary: A base-model result does not prove the released workflow.

02

What failure is unacceptable?

Use-case harms, reviewer rubric and exception paths. Accountable owner: Product owner with relevant risk specialists. Boundary: A high mean score can conceal a serious failure type.

03

What does the release result support?

Evaluation record, limitations and monitoring plan. Accountable owner: Release approver. Boundary: Evaluation does not certify compliance or eliminate risk.

07

Practitioner checklist

A working check before committing the next stage.

  • 01

    Version the complete workflow configuration, not only the model name.

  • 02

    Record evaluation-case provenance, approval and data handling boundary.

  • 03

    Include ordinary, difficult, uncertain and safe-refusal cases.

  • 04

    Choose measures and reviewer rubrics linked to the intended decision.

  • 05

    Set release criteria and unacceptable failure types before review.

  • 06

    Publish a bounded release and monitoring record.

Direct answers

Questions to settle before implementation

01Is a benchmark enough to release an AI feature?

No. Benchmarks can inform a decision, but a production release needs evidence about the actual workflow, users, inputs, controls and failure handling.

02How much evaluation is enough?

There is no universal number. The appropriate depth depends on impact, variability, data sensitivity, users, operational controls and the consequences of error.

Source discipline

Primary guidance and technical references

A practical next step

Turn the question into a scoped technical decision.

Bring the affected workflow, current system, constraints, decision owner and required evidence. We will assess the smallest responsible next step before proposing dates or delivery scope.

Start a technical conversation