Skip to main content
Start a conversation

Pilot planning tool

30-day AI pilot entry criteria

A pilot earns its place when it has real users, a defined decision, controlled inputs and a clear scale, iterate or stop gate.

When to use this

Use this after an opportunity or prototype has been identified and before calling it a 30-day pilot. It helps a team decide whether it is ready for controlled use.

  1. 01

    A named pilot owner and cohort

    Identify who is accountable for the pilot and which users will take part. A pilot without an owner or a defined user group cannot produce a reliable operating decision.

  2. 02

    A bounded operational question

    State the workflow, expected benefit, input, output and the point where a person decides or acts. Avoid treating broad adoption as the first experiment.

  3. 03

    Known data and integration boundaries

    Document the inputs, source systems, personal or confidential data, access rights, supplier dependencies and the route for handling poor or missing inputs.

  4. 04

    Usable human oversight

    Define who reviews an output, how they challenge or override it, what happens when the service is unavailable and how an incident or user concern reaches an owner.

  5. 05

    Feedback and measurement

    Choose what will be observed: completion, quality, exceptions, user feedback, safety concerns, cost or another project-specific measure. Do not use a pilot to manufacture an outcome claim.

  6. 06

    A close-out decision

    Set a review date and authority for scale, iterate, pause or stop. Record the assumptions and exclusions so that a pilot result is not overstated as general deployment evidence.

Applied guidance

Use the tool as a decision record, not a box-ticking exercise.

01

Frame the pilot as controlled learning

A pilot is most useful when it answers a specific operational question under known conditions. It is not a broad adoption announcement or a way to defer every design decision until later. Name the cohort, the workflow, the period, the inputs permitted and the condition under which a person should intervene. This lets users understand the purpose of the trial and lets stakeholders interpret the evidence honestly. A small, well-defined pilot that finds a limitation can be more valuable than a large trial that produces activity but no credible decision about whether to continue.

02

Choose users who can influence the work

Pilot participants should be close enough to the workflow to recognise whether the result is helpful, misleading or too slow. They also need a way to report issues and a reasonable understanding of what the system is and is not intended to do. A pilot owner should avoid selecting only enthusiastic early adopters if the future workflow will include different skill levels or operating conditions. Instead, document the cohort and the limitation of that choice. The result is not automatically representative of every user, but it is a clear basis for deciding which group should be tested next.

03

Use measurements that change a decision

Choose a limited set of measures because each will inform a named choice. Completion or turnaround time may show whether the workflow fits the job; sampled review may show whether outputs are usable; correction and escalation patterns may expose a weak boundary; cost or latency may affect operational viability. Capture the conditions around the result, including inputs, user experience and system availability. Avoid converting a short pilot into an unsupported productivity claim. The evidence should support an internal decision about the next controlled step, not a universal assertion about performance.

04

Build a real fallback

People must know what to do if the AI feature is unavailable, produces a doubtful result or operates outside its intended scope. The fallback may be the existing manual route, a draft that requires review, an exception queue or a pause control with a named owner. It must be usable in the moment, not merely described in a governance document. Test it during the pilot. A reliable fallback allows a team to learn without forcing users to choose between accepting an unsuitable output and stopping their work entirely.

05

Close the pilot with bounded conclusions

At the review date, compare the observed evidence to the original purpose and acceptance conditions. Record what worked, where human intervention was needed, which data or integration constraints appeared, and what the result does not establish. The next decision may be to expand to another cohort, iterate the workflow, add controls, pause the work or stop it. This avoids treating an early result as a complete production-readiness decision. It also preserves the evidence needed when new users, data sources or model configurations change the conditions that were actually tested.

06

Prepare users for the pilot boundary

Before the pilot begins, give participants a concise explanation of the intended task, the information they may provide, the limits of the output, the review expectation and the route for a concern. This is not a substitute for training or formal policy where those are required, but it prevents the system being treated as a general answer engine simply because it is available. The briefing should also identify the support contact and the approved workaround. Clear participant guidance makes feedback more useful because users can describe a failure against the workflow the pilot was designed to test.

Working prompts to adapt

Pilot purpose

The pilot tests whether [users] can use [workflow] with [AI capability] to support [specific decision], under [limits].

Control route

If the output is uncertain, unavailable or challenged, [role] will [review/override/escalate] using [evidence or process].

Close-out

On [date], [owner] will review [measures and evidence] and decide to [scale/iterate/pause/stop].

Primary guidance