What configuration is under evaluation?
Versioned model, prompt, retrieval, tools, UI and policy controls. Accountable owner: Technical/product owner. Boundary: A base-model result does not prove the released workflow.
Technical guide · Product, engineering, data and risk teams releasing AI-enabled workflows.
Production AI evaluation is the disciplined comparison of a defined system configuration against representative tasks, known failure modes and release criteria. It is not a single benchmark score, a successful demo or a claim that the model is generally reliable.
Working position
A model is rarely the whole system. Users encounter instructions, prompt templates, retrieval, business rules, tools, interface controls, permissions, human review and fallback behaviour. An evaluation that tests only a base model can be useful research, but it may not answer whether the released workflow works. Identify the versioned configuration: model or service, prompt, retrieval source, tool permissions, policy rules, UI path and environment. That is the thing for which a release decision is being made.
The target should be expressed in user and operational terms. What task is supported, what output is expected, what must never be automated, what evidence must be shown and what happens on uncertainty? This prevents the team from optimising an attractive metric while missing the decision that gives the product its value. It also makes later changes comparable rather than treating every model update as an entirely new unknown.
Delivery reasoning
Evaluation examples need traceable provenance. Record where each case came from, why it represents the intended workload, whether it includes personal or confidential information, who approved its use and how long it may be retained. Use a mix of ordinary cases, difficult cases, incomplete inputs, conflicting evidence, adversarial or misleading inputs where relevant and situations in which the correct answer is to decline, escalate or ask for more information.
Do not let the evaluation set become a hidden training or prompt-tuning loop without recording it. If examples influence instructions, retrieval or rules, separate development and hold-out cases where practical. Small teams do not need an elaborate laboratory to benefit from this discipline; they do need enough separation to avoid claiming that a system works simply because it was repeatedly tested against cases it was shaped around.
Delivery reasoning
Accuracy can matter, but it is often incomplete. A classification workflow may need precision and recall by class; a drafting workflow may need reviewer acceptance, traceability and harmful-claim rate; a retrieval workflow may need groundedness, citation usefulness, latency and abstention behaviour. Establish the unit of review and the reviewer rubric. Where expert judgement is required, record how reviewers are calibrated and how disagreement is handled rather than presenting one person's view as an objective benchmark.
Set release criteria before reading the results. They can be thresholds, required controls, unacceptable failure types or a qualitative decision record. The crucial point is that criteria need to be connected to the use case. A strong average can conceal an unacceptable failure in a high-impact route. NIST's AI RMF and playbook are voluntary frameworks, but their emphasis on governing, mapping, measuring and managing risks provides a useful structure for this decision.
Delivery reasoning
A credible evaluation makes failure visible. Test missing context, ambiguous requests, conflicting instructions, invalid tool inputs, retrieval with no useful result, content that should be refused, outages, slow responses and user attempts to bypass intended controls. If the product uses tools or downstream systems, verify permission boundaries and the result of a partial failure. The correct system behaviour may be a safe stop, a request for clarification or an escalation—not a fluent answer.
For workloads processing personal data or supporting decisions about people, involve the relevant privacy, legal, security or sector specialists early. Evaluation evidence cannot substitute for the underlying obligations. It can, however, make the system's capabilities and limitations more inspectable, which is necessary for a responsible release conversation.
Delivery reasoning
A useful release record names the configuration evaluated, cases and time period, measures, results, known limitations, approval decision and owner. It should state whether the result supports a limited release, further iteration or a stop. Attach evidence in a format the reviewers can inspect without exposing sensitive evaluation data unnecessarily. A short decision table is often more valuable than a large dashboard with no explanation of the data behind it.
Release is not the final measurement. Define what will be observed in use, how user feedback and exceptions are collected, what drift or regression signal matters, who reviews it and which actions are possible. If a model, prompt, retrieval source or tool behaviour changes, decide whether the change is minor, requires targeted regression checks or needs a new release review. The answer is workload-specific; the record makes that judgement repeatable.
Decision record
Versioned model, prompt, retrieval, tools, UI and policy controls. Accountable owner: Technical/product owner. Boundary: A base-model result does not prove the released workflow.
Use-case harms, reviewer rubric and exception paths. Accountable owner: Product owner with relevant risk specialists. Boundary: A high mean score can conceal a serious failure type.
Evaluation record, limitations and monitoring plan. Accountable owner: Release approver. Boundary: Evaluation does not certify compliance or eliminate risk.
Practitioner checklist
Version the complete workflow configuration, not only the model name.
Record evaluation-case provenance, approval and data handling boundary.
Include ordinary, difficult, uncertain and safe-refusal cases.
Choose measures and reviewer rubrics linked to the intended decision.
Set release criteria and unacceptable failure types before review.
Publish a bounded release and monitoring record.
Direct answers
No. Benchmarks can inform a decision, but a production release needs evidence about the actual workflow, users, inputs, controls and failure handling.
There is no universal number. The appropriate depth depends on impact, variability, data sensitivity, users, operational controls and the consequences of error.
Source discipline
A practical next step
Bring the affected workflow, current system, constraints, decision owner and required evidence. We will assess the smallest responsible next step before proposing dates or delivery scope.
Start a technical conversation