Skip to main content
Start a conversation

Responsible AI · 9 min read

The operating questions behind responsible AI

Responsible AI becomes practical when broad principles are translated into decisions that a product team and an operator can see, own and revisit.

01

Purpose and limits

State what the system is for, which users and people it affects, the decision it can support and the situations in which it must not be used. A narrow purpose makes evaluation, explanation and accountability possible. It is also the starting point for deciding whether a proposed use belongs in a product at all.

02

Inputs, data and authority

Document the information that enters the workflow, where it comes from, who can access it, what the system returns and which person or role has authority to act. This is not a generic data-residency or compliance promise. It is a project-specific decision about the data, provider, environment, retention, access and safeguards actually selected.

03

Human review that can be used

Define when a person reviews, overrides or stops the system, then make that path usable in the interface and operating process. A reviewer needs the relevant context, a meaningful ability to challenge an output and a way to record an exception. Human review written only in a policy document is not a dependable product control.

04

Evaluation before scale

Agree what should be checked before the system is relied on more widely. Depending on the use case this may include relevance, accuracy, robustness, safety, latency, cost, bias, user experience or operational impact. The test should include the difficult and uncertain cases, not only examples selected because they work well.

05

Evidence and change over time

Record material model, data, prompt, configuration, control and performance decisions so later changes can be understood. Monitoring should lead to an owner and an action: investigate, adjust, pause, revert or re-evaluate. This creates an operating record, not a claim that every risk has been eliminated or that a system is certified.

06

Start with the impact of being wrong

The relevant question is not whether a system is called AI. It is what happens if the output is irrelevant, incomplete, biased, unavailable, disclosed to the wrong person or acted on too confidently. Identify the people affected, the operational consequence, the existing fallback and the person who can decide that a result should not be used. A low-impact drafting aid and a workflow that influences eligibility, care, employment, finance or safety need different boundaries and evidence. Naming the impact early helps the team decide whether assistance is appropriate at all and what review route is proportionate.

07

Turn governance into observable controls

A policy can state that a human remains in the loop while the product offers no usable way to inspect, challenge or stop an output. Convert the principle into behaviour: what information is shown, what confidence or limitation is disclosed, when the workflow pauses, who may override it, how an exception is recorded and what happens next. The control should work under ordinary operational pressure, not only during a demonstration. If the only route is for a user to remember a separate policy or send an unstructured email, the system has not yet made oversight operational.

08

Evaluate representative failure, not a curated demo

Evaluation should include the ordinary cases, but it should also include missing context, ambiguous input, changed source material, adversarial or unexpected content and cases where a reasonable person should disagree with the proposed output. The test set, acceptance threshold and reviewer should be appropriate to the use case. Do not translate a promising result into a claim of accuracy, fairness, safety or compliance without evidence that supports that exact claim. Record what the evaluation does not cover as well as what it does, because the boundary is often the most useful information for the next operating decision.

09

Review changes as changes to the system

A model setting, prompt, retrieval source, provider endpoint, access permission, data source or workflow rule can materially change how the system behaves. Treat it as a change that needs an owner, a reason, a record and, where relevant, a renewed evaluation. This does not require every experiment to become a lengthy approval process. It requires the team to recognise when an apparently small change alters the purpose, data boundary, user experience or risk. A reversible pilot is valuable only if the team can tell what changed and can pause, revert or redesign when the evidence calls for it.

010

Give incidents and concerns a usable route

People need to know how to report a harmful, unsuitable or unexpected result and what will happen after they do. Define the receiving role, the information needed for investigation, how affected records are protected and when use should be restricted or paused. Link the route to the wider incident, privacy, security or safeguarding process that applies to the organisation rather than creating an isolated AI inbox. During a pilot, exercise the route with a representative scenario. The objective is not to claim that incidents cannot occur; it is to ensure that a concern can become an owned investigation, a bounded decision and an evidenced change.

Have an urgent brief behind the question?

Discuss an urgent project