Why an AI project starts with the process, not the model
Most failed AI projects break not because the model was weak but because the process was never described.
A failed AI project rarely looks like a technical accident. More often it looks like this: everyone liked the demo, the pilot went live, and two months later nobody uses it. The post-mortem almost always lands in the same place — nobody described in advance at which point of the working day, and in place of which action, a person is supposed to turn to the AI.
The model has almost nothing to do with it. It answers exactly as asked, on the data it was given. Everything else is process.
Roles and actions
The first thing to write down is who takes part and what they do. Not “support”, but: the first-line operator receives an enquiry, looks for an answer, replies to the client and calls a senior colleague when in doubt. Each of those is a separate point, and AI plugs into one of them — not into “support” as a whole.
Once the actions are written out, the bottleneck usually becomes obvious. Sometimes it is not where you expected: the time goes not into wording the answer but into finding the current version of a document. Then it is a different task altogether.
It is also worth writing down what happens after the action. Who checks the result, where it goes, what happens on a mistake and who notices it. Half the architectural decisions about a scenario are made right here: if nobody would see the mistake, it is too early to automate that action, however convenient it looks.
Data sources
AI answers from what it was given. So before launch you name the specific sources: these policies, this section of the knowledge base, these CRM records. Not “all our documents” — there are usually more of them than anyone thinks, and half are out of date.
- each source has an owner who keeps it current
- it is clear how often it changes and what happens when it does
- it is visible which data is restricted, and from whom
- there is a decision on what to do when sources contradict each other
Limits
Limits are not a list of prohibitions for safety’s sake but a description of the area where the scenario makes sense at all. What the AI answers, what it stays silent about, what it does when the data is missing, and who it hands the question to when it falls outside.
The “I don’t know” rule is worth building in from the start. “That is not in the documents” is a useful answer. A plausible invention in its place costs more than no answer at all.
Limits are easier to state through examples than through rules. Collect a dozen real enquiries the scenario must handle, and a dozen it must not. The line between those two piles is the area of use — and it usually turns out narrower than anyone assumed at the start.
The boundaries are worth agreeing not only with the team but with a lawyer or whoever owns risk — especially when the scenario touches money, commitments to clients or personal data. The conversation takes an hour and spares you the situation where a finished scenario gets switched off just before launch for a reason nobody thought about.
Answer quality
“It answers well” is not a criterion. A criterion looks different: on a sample of real enquiries the operator accepted the prompt in so many cases, and in the rest had to search by hand. That number can be measured before launch and repeated a month later.
It helps to assemble a set of questions the scenario is always checked against: typical, rare, contentious and deliberately out of scope. That same set becomes the test after every change to the prompt or the document base.
It matters that the set is assembled not by the developer but by whoever does the work today. A developer will invent questions the system can already answer; an operator brings the ones they stumble over themselves. The second kind is more useful: those are what show where the scenario will meet reality.
A safe MVP
A safe first step is simple: the AI prompts, a human decides, and everything is visible in the logs. No actions in external systems, nothing sent to the client directly. At that level a mistake costs a few seconds, and enough data for judgement accumulates fast.
Widening the scenario makes sense only once the statistics are in: where the AI helps, where it gets in the way, which questions it consistently fails. Those same observations usually suggest the next scenario, too.
One more thing: an MVP needs a way to be switched off, described in advance. Not because we expect it to fail, but because without one the team keeps using the scenario out of habit even when it gets in the way. An agreement that “if the number is not there in two months, we switch it off and look again” saves far more than it sounds.
Want your process unpacked before picking the technology?
Describe who takes part, which actions repeat and where the data lives. We will work out which scenario is safe here and what needs preparing before launch.
Show us your process