Skip to content
Rivl
25 August 2026AI9 min

How to use AI in operations without breaking the business

The demos are all about replacing a person. The wins are all about removing a step nobody wanted to do in the first place.

Most operations teams have now had the same year. Somebody trialled a chat assistant, it produced something impressive in the demo, and then it turned out that the ten minutes it saved were spent checking whether the output was right. The trial ends, everyone concludes the technology is not ready, and the actual opportunity goes unexamined.

The opportunity is real, and it looks nothing like the demo. It is narrow, unglamorous, and usually invisible to whoever approved the budget.

The sections of this note, from the failure test through to the honest limit.
How this note is structured, from the failure test to the honest limit

The test that separates the two

One question decides whether an AI step will pay: what happens when it is wrong. Not how often, what happens. If a wrong output is caught immediately and costs a second to fix, the step is a good candidate even at mediocre accuracy. If a wrong output flows downstream into an invoice, a shipment or a customer's inbox, then accuracy has to be near perfect before it saves anything at all, and near perfect is expensive.

This is why summarising a call transcript works and why automatically approving a supplier invoice does not. Both are the same technology. Only one of them fails safely.

Where it genuinely pays in operations

TaskWhy it worksWhat it does not do
Drafting a first replyHuman edits before sendingReplace the human review
Extracting fields from documentsErrors visible against the sourceRemove the need for the source
Categorising incoming requestsWrong category is cheap to fixDecide priority on its own
Searching your own past decisionsAnswer is checkableBe authoritative about policy
Turning notes into a structured recordWrong entry is edited in placeReplace the note taker
Operational tasks where an AI step pays, why each one works, and the thing it does not remove.
Where an AI step pays in operations, and what it does not remove

Read down the middle column and the pattern is obvious. Every entry works because a person sees the output next to the thing it came from, before it matters. That is the whole design principle, and it is more useful than any list of tools.

Where it quietly costs more than it saves

Three patterns account for most of the disappointment we see.

The verification tax

If checking the output takes as long as producing it, you have moved work rather than removed it, and you have made it duller. Watch for this in anything where the person now reads carefully to catch a plausible mistake. Plausible wrong answers are more expensive than obviously wrong ones.

Automating a step that should not exist

The most common failure is not technical. A team automates the assembly of a weekly report that nobody has acted on in a year. The step gets faster and the waste becomes permanent, because now it is cheap enough that nobody questions it. Before automating anything, ask what decision the output drives, which is the same discipline as any reporting build and the reason automating manual reporting without buying a BI platform starts with the decision rather than the tool.

Paying for live when current is fine

AI features are often sold attached to real-time infrastructure that the use case does not need. The distinction between genuinely live and merely current is the same one we work through in what real-time actually means for a sales dashboard, and it usually saves more money than the AI saves.

The governance part, briefly and without drama

You do not need a policy document to start. You do need to be able to answer three questions before an AI step touches anything that leaves the building: what data goes into it, who checks the output, and what happens when it is wrong.

If you want a structure rather than an ad hoc answer, the NIST AI Risk Management Framework, released in January 2023, is the most usable public reference we know of. It is explicitly intended for voluntary use, and it organises the work into four functions: Govern, Map, Measure and Manage. For a small operations team that reduces to something practical: decide who owns it, write down where it is used, agree how you will know it is working, and review it. That is a morning's work, not a programme.

The reason to do it at all is not compliance theatre. It is that AI steps spread sideways. One person's useful shortcut becomes six people's undocumented dependency, and nobody can say what the system does any more.

How to run a trial that tells you something

  • Pick one task, not one department.
  • Measure the current version first, in minutes, over a real week.
  • Include the checking time in the after measurement. This is the step everyone skips and it is the step that decides the answer.
  • Set the stopping rule before you start, so ending it is not a failure.
  • Keep the manual path working throughout.

That last point matters more than it sounds. The trials that end badly are the ones where the old process was switched off in week one and there was nothing to fall back to in week three.

The honest limit

Most operational pain in small and mid-sized companies is not an AI problem. It is that the data lives in four systems that do not talk to each other, and no model fixes that. If your team spends its day re-keying the same information between a spreadsheet, a CRM and an accounting package, the return on connecting those three is larger and more certain than anything a model will do on top of them, as the case in when a stock spreadsheet starts costing more than it saves shows.

Marketing teams tend to hit this earlier than operations does, because their reporting is spread across more platforms. KF Agency's Arabic guide to automating marketing reports with AI (written in Arabic) covers that side of it in more detail than we do here.

Used narrowly, on tasks that fail safely, this is one of the better returns available to an operations team right now. Used as a replacement for deciding how the work should run, it is an expensive way to keep the current mess at higher speed.

Describe it. We build it.

Seven or twelve days, pay on delivery, a year of maintenance included. Bring the problem, not a spec.

Book a meeting

Read next