What AI agents can actually do (and what they can't yet)

On a benchmark of real business workflows, the leading AI models finish well under half of them unaided. Here's what that means if you're thinking of handing work to an agent.

A small blue robot holding up a lit lightbulb

There’s a particular kind of sales pitch going round at the moment. It says AI agents can now run your operations for you — take the enquiry, check the system, raise the order, chase the payment, no human needed.

We build AI agents for a living, so you’d expect us to nod along. We’re not going to, because there’s a number that makes the pitch difficult to defend, and it comes from the companies making the models.

The number

In April 2026, Zapier published AutomationBench — a test of whether AI models can complete realistic end-to-end business workflows. Not puzzles or exams. Tasks across sales, marketing, operations, support, finance and HR, each one dropping the model into an environment with, in Zapier’s words, “a CRM with live data, an inbox with threads, a calendar with conflicts… similarly named contacts, inconsistent formats, multi-step tool chains where a wrong call cascades.”

The scoring matters. It’s deterministic — the final state of the system is checked against what should have happened. There’s no AI grading another AI’s homework. Zapier is well placed to build this: 3.7 million companies use it, and it processes around two billion AI tasks a month.

Then in July, OpenAI published its GPT-5.6 results and included an AutomationBench table. Its own flagship, GPT-5.6 Sol, scored 18.1%. Its mid-tier Terra scored 15.2%. Anthropic’s Claude Fable 5 scored 17.4%. Google’s Gemini 3.5 Flash, 14.5%.

Read that again. On realistic, ambiguous, multi-system business workflows, the best models available complete fewer than one in five without help. And that table is published by OpenAI, on a third party’s benchmark. It’s not a critic’s number.

That table is three weeks old, and things move. Anthropic’s Claude Opus 5, released 24 July 2026, claims a pass rate “around 1.5× the next-best model for the same cost per task” on the same benchmark. It’s a vendor claim about its own product, so treat it accordingly — but take it entirely at face value and you land somewhere around a quarter to a third. Better. Still nowhere near “leave it to run”.

So why does everyone keep saying agents work?

Because for certain kinds of task, they genuinely do, and the improvement over the last eighteen months has been startling.

The clearest example is computer use — an AI operating software by looking at the screen and using a mouse and keyboard, the way a person does. In March 2026, OpenAI’s GPT-5.4 scored 75.0% on OSWorld-Verified, which measures exactly that: navigating a desktop through screenshots and keyboard and mouse actions. The human baseline OpenAI cites alongside it is 72.4%.

Worth a pinch of salt — that human figure is carried over from the original 2024 research paper rather than measured against the model on the day, and OpenAI ran its evaluation at maximum reasoning effort in a research setting. But the shape of it is real, and it was unthinkable eighteen months ago.

That matters enormously for British businesses, because so much of what you deal with has no API and never will. Council portals. Insurer extranets. The supplier ordering system from 2011. The bit of software your industry runs on that has four hundred customers and no integration story. Until recently, automating any of that meant fragile screen-scraping. Now it’s a solvable problem.

The gap between “beats humans at driving a computer” and “finishes 18% of real business workflows” is the whole ballgame. Individual steps are close to solved. Stringing thirty of them together, across four systems, when two contacts have the same surname and the date format changed halfway through 2023 — that’s not solved.

What the people deploying them have learned

Zapier surveyed 525 senior people at large American companies in late 2025. Two caveats before the numbers: it’s a vendor survey, and it covers US firms with 1,000 or more employees, so it’s not a read on a Yorkshire SME. The pattern is still instructive.

The departments most likely to have agents in production were customer support (49%) and operations (47%). Finance was last, at 24% — which tells you people put agents where mistakes are recoverable and keep them away from the money.

And the headline finding on how they’re managed: human-in-the-loop was the most common approach, at 38%. Only 20% ran systems autonomously with minimal oversight. The organisations furthest ahead are, overwhelmingly, not letting the things run unsupervised.

DSIT’s UK research, from 2025 fieldwork, found agentic AI was the least-adopted technology among British AI users — 7% of them. We are very early here.

The failure rate nobody quotes at you

In May 2026, Gartner published a prediction that deserves more attention than it got: by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps found only after something went wrong in production.

Note demote or decommission. Not “fail to build” — build, deploy, then pull back. Gartner’s diagnosis is that firms treat agent governance as a switch: either locked down, which slows everything and pushes people into building their own unofficial versions, or trusted, which is fine right up until it isn’t.

Gartner’s four-level framework is the most useful thing we’ve read on this all year, and it’s plain enough to use in a meeting:

  1. Observe — the agent reads and reports. It can’t change anything.
  2. Advise — it drafts and recommends. A person carries it out.
  3. Act with approval — it can write, send or modify, but every action needs a sign-off.
  4. Act autonomously — it acts within guardrails; a human reviews exceptions.

Our own observation, for what it’s worth: most vendors are selling level four. Most of the deployments we see working are sitting at two or three — which is roughly what Zapier’s survey found too, with human-in-the-loop the most common approach and only a fifth running with minimal oversight.

Gartner’s warning about level three is the sharpest line in the release, and it matches what we see: approvals “can degrade under time pressure or approval fatigue, creating a false sense of safety while expanding the attack surface”. If a person has to click Approve four hundred times a day, by Thursday they aren’t reading. You’ve built an autonomous agent with extra steps and a false paper trail.

One more, for the sceptics who want a bigger number: Gartner predicted in June 2025 that over 40% of agentic AI projects would be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. That’s still the most-quoted failure statistic going. It’s from 2025, and it’s worth dating it properly when someone waves it at you.

What this means if you’re considering it

None of the above is an argument against automating things. It’s an argument about where to point it.

Start at Advise, not Autonomous. Have the agent do the reading, the matching and the drafting, and have a person press send. You get most of the time saving and none of the risk. Move to approval-based action once you’ve watched it be right for a few weeks and you know what its mistakes look like.

Pick jobs where being wrong is visible and cheap. Drafting a reply someone checks: good. Categorising expenses a bookkeeper reviews: good. Silently adjusting stock levels or issuing credit notes: not yet.

Count the steps. The 18% figure is about long chains across multiple systems. A three-step job in two systems is a very different proposition to a thirty-step job across five, and most of the value in a small business is in the short chains anyway.

Design the approval so it stays real. If sign-off is going to happen hundreds of times a day, it will stop being sign-off. Better to auto-handle the clear cases and route only genuine exceptions to a person — fewer decisions, each one actually made.

The honest summary is this: the tools got dramatically better at doing single things, and only modestly better at doing whole jobs. The businesses getting value are the ones who worked out which of their jobs are actually a short chain of single things — and who kept a human where the consequences are.

Get in touch

Come and have a natter

Whichever suits you best — we're not fussy.

Or email us straight off: eyup@ai-up.co.uk