Skip to main content
Start a conversation

Responsible AI tool

Responsible AI delivery evidence pack

Responsible AI becomes operational when a team can show what the system is for, who owns it, how it is evaluated and what happens when it is uncertain or wrong.

When to use this

Use this to prepare the evidence a product, delivery or governance team needs before a controlled AI release or pilot. Make the pack proportionate to the use case and impact.

  1. 01

    Purpose and prohibited use

    State the intended purpose, users, people affected, decision supported and explicit limits. A clear limit is necessary for meaningful evaluation and accountability.

  2. 02

    Inputs and data decisions

    Record data sources, permissions, quality assumptions, retention, access, supplier settings and known gaps. Link the record to project-specific privacy and security review where applicable.

  3. 03

    Human authority and escalation

    Show when a person reviews, overrides or stops the workflow; what context they receive; and how a concern, incident or harmful output reaches an accountable owner.

  4. 04

    Evaluation plan

    Choose the relevant tests for the purpose: relevance, accuracy, robustness, safety, latency, cost, user experience, fairness or operational effect. Include difficult, uncertain and failure cases.

  5. 05

    Release and change record

    Record the model, prompt, configuration, integration, version, limitations, acceptance criteria and approval for the release. Keep later changes traceable to their reason and evidence.

  6. 06

    Monitoring and action

    Define what will be observed, who reviews it, how often, and the actions available: investigate, adjust, pause, revert, retrain or retire. Monitoring without an owner is not an operating control.

Applied guidance

Use the tool as a decision record, not a box-ticking exercise.

01

Describe the complete configured system

The evidence pack should identify more than a model name. It should describe the user task, input boundary, system instructions, retrieval or context sources, tools or APIs, permissions, interface, human review route and the output that may influence work. This makes it possible to evaluate the product that will actually be used rather than an abstract base model. If a component changes later, the team can see whether it alters the previously assessed behaviour. The record can be concise, but it should be specific enough for an accountable owner to understand what is inside and outside the release boundary.

02

Use representative and difficult evaluation cases

Evaluation should test the work users are expected to perform, including incomplete inputs, conflicting information, requests outside intended use, tool failures and cases where the appropriate response is to ask for clarification, defer or refuse. Agree what reviewers are looking for: task usefulness, factual support where relevant, safe tool behaviour, escalation and user comprehension. A score can be useful, but it should not hide important failure types. Retain the configuration reference, case set and review outcome so that later changes can be compared to the evidence that informed the initial release decision.

03

Make human oversight usable

Human oversight is meaningful only when a person has enough context, authority and time to act. The record should state which outputs require review, what evidence the reviewer sees, how they correct or override the result, and how a concern is escalated. It should also explain what happens when the service is unavailable or when a user identifies a harmful or unsuitable output. A checkbox that says ‘human in the loop’ is not sufficient. The workflow needs to show the actual moment at which authority remains with a person and how that intervention is recorded.

04

Connect monitoring to actions

After release, collect only the signals that support a real operating decision. Depending on the workflow, these may include errors, latency, user corrections, exceptions, sampled output review, feedback, tool failures, configuration changes or cost. Assign a reviewer, a review cadence and the actions available: investigate, adjust instructions, restrict a tool, pause a feature, revert configuration or retire the use case. Logging everything by default can create its own privacy and security burden. Monitoring should be proportionate to the purpose, data sensitivity and impact of the system being operated.

05

Keep the pack as a living decision record

The first release is not the end of the evidence story. New data sources, user groups, prompts, model versions, retrieval material, tools or policies can all alter the behaviour that was evaluated. Update the pack when a meaningful change is proposed, record whether targeted re-evaluation is needed, and preserve the rationale for the decision. This does not confer legal compliance, certification or a guarantee of safe outcomes. It gives the organisation a practical record for explaining how it chose to build, release, monitor and improve a defined AI-enabled workflow.

06

Keep claims proportionate to the evidence

The pack should help a team avoid turning responsible delivery activity into unsupported public claims. A review of representative cases is evidence for the cases and configuration reviewed; it is not proof that every output will be correct. A human oversight step is a workflow control only if it can be exercised in practice. A security or privacy review may identify actions and conditions but does not by itself certify the system. State the scope, date and owner of the evidence. This improves decision quality and gives customers, users and internal reviewers a clearer picture of what the release can reasonably support.

Working prompts to adapt

Purpose statement

This system supports [role] to [task/decision] using [inputs]. It must not be used for [prohibited use or decision].

Oversight statement

[Role] reviews [condition/output] and can [override/stop/escalate]. The evidence available to them is [context].

Release record

Version [identifier] was evaluated against [criteria/cases], with known limitations [limits]. [Owner] approved/rejected the next release decision on [date].

Primary guidance