How to use AI in operations without breaking the business
The demos are all about replacing a person. The wins are all about removing a step nobody wanted to do in the first place.
Most operations teams have now had the same year. Somebody trialled a chat assistant, it produced something impressive in the demo, and then it turned out that the ten minutes it saved were spent checking whether the output was right. The trial ends, everyone concludes the technology is not ready, and the actual opportunity goes unexamined.
The opportunity is real, and it looks nothing like the demo. It is narrow, unglamorous, and usually invisible to whoever approved the budget.
The test that separates the two
One question decides whether an AI step will pay: what happens when it is wrong. Not how often, what happens. If a wrong output is caught immediately and costs a second to fix, the step is a good candidate even at mediocre accuracy. If a wrong output flows downstream into an invoice, a shipment or a customer's inbox, then accuracy has to be near perfect before it saves anything at all, and near perfect is expensive.
This is why summarising a call transcript works and why automatically approving a supplier invoice does not. Both are the same technology. Only one of them fails safely.
Where it genuinely pays in operations
| Task | Why it works | What it does not do |
|---|---|---|
| Drafting a first reply | Human edits before sending | Replace the human review |
| Extracting fields from documents | Errors visible against the source | Remove the need for the source |
| Categorising incoming requests | Wrong category is cheap to fix | Decide priority on its own |
| Searching your own past decisions | Answer is checkable | Be authoritative about policy |
| Turning notes into a structured record | Wrong entry is edited in place | Replace the note taker |
Read down the middle column and the pattern is obvious. Every entry works because a person sees the output next to the thing it came from, before it matters. That is the whole design principle, and it is more useful than any list of tools.
Where it quietly costs more than it saves
Three patterns account for most of the disappointment we see.
The verification tax
If checking the output takes as long as producing it, you have moved work rather than removed it, and you have made it duller. Watch for this in anything where the person now reads carefully to catch a plausible mistake. Plausible wrong answers are more expensive than obviously wrong ones.
Automating a step that should not exist
The most common failure is not technical. A team automates the assembly of a weekly report that nobody has acted on in a year. The step gets faster and the waste becomes permanent, because now it is cheap enough that nobody questions it. Before automating anything, ask what decision the output drives, which is the same discipline as any reporting build and the reason automating manual reporting without buying a BI platform starts with the decision rather than the tool.
Paying for live when current is fine
AI features are often sold attached to real-time infrastructure that the use case does not need. The distinction between genuinely live and merely current is the same one we work through in what real-time actually means for a sales dashboard, and it usually saves more money than the AI saves.
The governance part, briefly and without drama
You do not need a policy document to start. You do need to be able to answer three questions before an AI step touches anything that leaves the building: what data goes into it, who checks the output, and what happens when it is wrong.
If you want a structure rather than an ad hoc answer, the NIST AI Risk Management Framework, released in January 2023, is the most usable public reference we know of. It is explicitly intended for voluntary use, and it organises the work into four functions: Govern, Map, Measure and Manage. For a small operations team that reduces to something practical: decide who owns it, write down where it is used, agree how you will know it is working, and review it. That is a morning's work, not a programme.
The reason to do it at all is not compliance theatre. It is that AI steps spread sideways. One person's useful shortcut becomes six people's undocumented dependency, and nobody can say what the system does any more.
How to run a trial that tells you something
- Pick one task, not one department.
- Measure the current version first, in minutes, over a real week.
- Include the checking time in the after measurement. This is the step everyone skips and it is the step that decides the answer.
- Set the stopping rule before you start, so ending it is not a failure.
- Keep the manual path working throughout.
That last point matters more than it sounds. The trials that end badly are the ones where the old process was switched off in week one and there was nothing to fall back to in week three.
The honest limit
Most operational pain in small and mid-sized companies is not an AI problem. It is that the data lives in four systems that do not talk to each other, and no model fixes that. If your team spends its day re-keying the same information between a spreadsheet, a CRM and an accounting package, the return on connecting those three is larger and more certain than anything a model will do on top of them, as the case in when a stock spreadsheet starts costing more than it saves shows.
Marketing teams tend to hit this earlier than operations does, because their reporting is spread across more platforms. KF Agency's Arabic guide to automating marketing reports with AI (written in Arabic) covers that side of it in more detail than we do here.
Used narrowly, on tasks that fail safely, this is one of the better returns available to an operations team right now. Used as a replacement for deciding how the work should run, it is an expensive way to keep the current mess at higher speed.
Describe it. We build it.
Seven or twelve days, pay on delivery, a year of maintenance included. Bring the problem, not a spec.
Book a meetingRead next
How to choose a software development company
Portfolios and day rates tell you almost nothing. Seven checks that separate a firm that will finish from one that will hand you a repository nobody can maintain.
ReadA real-time dashboard for a sales team: what real means
Most teams asking for real-time need current, not live, and the gap is a large bill. What the word costs, what a five-minute refresh buys, and when live pays.
Read