Back to journal

What a useful AI pilot should prove.

A pilot earns its keep when it helps you make a decision. A polished demonstration is only one part of that evidence.

Set a small, meaningful boundary.

Choose one workflow, a defined group of users, and a realistic set of inputs. Keep the scope narrow enough to learn something specific, while preserving the parts that make the real work difficult.

A support pilot, for example, could focus on drafting responses for one category of ticket. That gives you a place to assess the quality of the draft, the context it uses, and the review it requires.

Agree on the decision in advance.

Before building, decide what would justify continuing. You might compare completion time, correction rates, source accuracy, or the number of steps the user still needs to perform.

Include the effort required to operate the solution. A faster draft may offer little value if checking it takes longer than writing from scratch.

Make room for ordinary failures.

Try incomplete inputs, conflicting documents, unfamiliar requests, and questions the system cannot answer. Watch whether the user can recognize and recover from a problem.

Human review needs a practical design: what the person can see, what they can change, and which decisions remain theirs. That should be part of the pilot itself.

End with an honest next step.

A pilot can lead to a production build, a revised approach, or a decision to stop. Each is a useful result when it follows from evidence.

The final handoff should explain what worked, what remains uncertain, and what production would require across integrations, access, monitoring, and ownership.

What could work better?

Let’s find out together