Notes on scoping an AI engagement

Journal Careers Home

The projects that go badly rarely go badly for technical reasons. They go badly because the first two weeks were spent agreeing on a direction rather than on a question, and a direction cannot be wrong in any way that anyone notices until a quarter has passed.

What follows is the short list we work through before agreeing to start. It is not a methodology. It is a set of things that, left unsettled, reliably cause trouble later.

State the decision the work informs

Not the deliverable — the decision. “A retrieval system over our documentation” is a deliverable. “Whether we can cut first-response time on support tickets enough to stop growing the team linearly with volume” is a decision. The second tells you what to measure, what precision is required, and when you are finished. The first tells you none of those things, which is why it is the more comfortable one to write down.

If you cannot say what would count as an answer, the engagement is not ready to start.

Find out what a competent baseline scores

Before any model work, establish what the obvious approach achieves. Sometimes that is a keyword search. Sometimes it is a hundred lines of rules written by the person who has done the task manually for three years. Sometimes it is the existing process, measured properly for the first time.

Two things come out of this. You get a number that every later result has to beat to mean anything. And occasionally you discover the baseline is already sufficient, at which point the most valuable possible outcome is to stop. We have had engagements end here. It is not a failure; it is the cheapest correct answer anyone got that year.

Settle what the data actually permits

The question is not whether data exists. It is whether the data that exists supports the thing being asked of it. Specifically: does it cover the cases that matter, or only the common ones? Is it labelled by someone whose judgement you would defend? Can it legally and practically leave the systems it currently lives in?

That last one is worth resolving in week one rather than week nine. A plan that depends on data which cannot move is not a plan, and discovering this late converts a scoping problem into a schedule problem.

Name the failure you cannot afford

Every system will be wrong sometimes. The useful question is which kind of wrong is tolerable. A system that occasionally declines to answer is a different engineering problem from one that must always answer and may occasionally be confidently incorrect. These require different architectures, different thresholds, and different amounts of human review.

Teams often leave this implicit, then discover during rollout that their tolerance was much narrower than the system was built for. Ask it at the start, get it in writing, and design toward it.

Decide who owns it afterwards

An external engagement that produces something nobody internally can maintain has created a liability with a nice demo attached. Identify the person or team who will hold it, involve them from the beginning, and treat the handover as part of the build rather than an event at the end.

This has a useful side effect. Work that has to be explained to a named person as it is built tends to come out simpler, because complexity that cannot be justified out loud rarely survives the conversation.

A note on timelines

Scoping done properly takes one to two weeks and feels like a delay. It is the cheapest fortnight in the project. The alternative is to spend it anyway, distributed across six months, in the form of rework.

All entriesNewerEvaluation is the product12 August 2026