← All articles
Operations • Field note

Your agent will be judged on numbers you never wrote down

Gartner expects more than four in ten agentic AI projects to be cancelled by the end of 2027, largely because nobody can show what they changed. The fix is a couple of weeks of counting, and it only works if you do it before launch.

David Soden • 5 min read • 18 September 2026
A hand marking a wooden plank with a pencil against a yellow tape measure
The mark goes on the board before the saw comes out. Afterwards there is nothing left to measure.

What happened

Mining Weekly ran a piece arguing that agentic AI is becoming a new layer in the business, sitting between people, data, applications and workflows. The distinction it draws is a useful one. A chatbot answers questions. An agent works inside a process: it reads the information, calls the tools it has been approved to use, and takes defined actions in your systems.

Its advice on where to start is ordinary and correct. Pick a process you already understand, such as IT requests, customer service, invoice processing or order management, where you can count the volume of work and time each task. It cites Gartner's prediction that over 40% of agentic AI projects will be cancelled by 2027 because of rising costs, unclear value and weak risk controls. And it says success depends on showing improvement over a clear starting point.

Why this matters to your business

That last line is the one most teams skip. The pilot gets approved on enthusiasm, the agent goes live, and three months later finance asks what it saved. Nobody wrote down how long a refund took before, how many arrived each week, or how often one had to be redone. So the answer is a handful of anecdotes, and anecdotes lose budget reviews.

The cost of skipping is lopsided. Measuring before launch takes a couple of weeks of counting work your team already does. Measuring after launch can't be done at all, because the process you would have measured has changed. The agent changed it, and you can't go back and time the old way.

A missing baseline also hides whether the agent is any good. An agent that finishes 60% of cases and hands the rest to a person could be a win or a loss, depending entirely on what those people were spending their time on before.

An antique silver stopwatch on a red and white cord lying on a wooden table
Time per case is the number almost nobody has on the day the agent launches.

Four numbers to write down before the agent starts:

Why this is a CX-Builder use case

The Mining Weekly piece lists what an agent needs before it can run a process: reliable data, access to the systems, a defined outcome, someone accountable, a path for exceptions, and monitoring. Most of that list is also what makes a baseline possible. A defined outcome tells you what to count. The exception path shows where the agent's limits are. Monitoring gives you the after number to put beside the before.

CX-Builder is built around that shape. A flow has a start and an end you can see on the canvas, so one finished case has a plain meaning. A human-in-the-loop gate marks the exception path, the point where the agent stops and a person decides, and every stop there is a record you can count. Because it runs on your own infrastructure, the run history sits with the rest of your data, where your team can compare it with the pre-launch numbers without asking a vendor for an export.

It also lets you start narrow, which is what the article recommends. An agent that drafts the invoice match and waits for approval is a smaller step than one that pays the invoice, and it is easier to measure. You count how many drafts were accepted without edits.

Two sheets of paper marked paid and due in red ink beside a calculator and a pair of glasses on a white desk
Invoice processing is on the article's shortlist because every item already has a clear finished state.

What this looks like if you build it

Build one flow for one process and run it alongside the team first. It takes live cases and writes its proposed answer to a log while people keep doing the work the old way, so two weeks later you have both numbers from the same cases. Then switch it on with an approval node ahead of any action that moves money or messages a customer, and track the share of proposals that go through unedited. When that share holds steady, raise the threshold on the gate rather than removing it.

Keep the escalations in view as well. Cases the agent hands back are the exception path doing its job, and their count and reasons tell you where the next round of work is.

A woman in a cap and glasses reading a padded envelope at a desk in a stockroom lined with boxes
The cases the agent sets aside for a person belong in the result.

The article's other figures show how early most businesses are. A South African survey it cites found 43% of companies still exploring AI and 13% using it in core processes. For most readers the first agent is still ahead of them, and that first pilot is the one worth measuring properly, because it sets how every later one gets judged.

The takeaway

Pick the process you plan to hand to an agent and spend the next two weeks counting it: weekly volume, time per case, rework rate, and what gets escalated. Put those four numbers in the project brief before anyone builds anything. If you can't get them, the process isn't ready for an agent yet.

All articles Install CX-Builder View on GitHub