Not a chatbot answering a question — an agent completing a task. Looking something up, updating a record, deciding what happens next, and moving between systems to finish the job.
An enquiry arrives written in a way the form didn’t anticipate. An invoice doesn’t match the purchase order for a reason a human would spot immediately. A supplier email says “we can do Thursday instead” and something has to work out what that means for three other bookings.
Those are the tasks that stay manual — not because they’re difficult, but because they need judgement applied to unstructured information. And they’re often the ones eating the most time.
That’s the gap agents fill. They’re also the newest and least proven part of this field, and this page is going to be honest about that.
This area is heavily oversold. Here’s our actual read, and we’ll update it as things change.
Reading unstructured input — an email, a document, a form filled in oddly — extracting what matters, and putting it in the right place. This is the strongest current use case.
Look up the customer, check their history, update the record, notify the right person, create the follow-up task. Each step is verifiable and the boundaries are clear.
Reading an enquiry, working out what it actually concerns, and sending it to the right place with a summary attached.
Producing the first version of a reply, a quote, a report — with a person checking it before it goes anywhere.
Watching for things that don’t look right and escalating them, rather than acting on them.
Anything financial, legal, contractual or safety-related needs a human checkpoint. Not because agents always get it wrong, but because they get it wrong occasionally and confidently, and you can’t tell which time is which without looking.
If a mistake surfaces immediately, an agent is fine — you’ll catch it. If a mistake sits undetected for three months and then costs you a client, don’t automate it without review.
When something goes wrong, someone has to be answerable. “The agent decided” is not a position you want to be in with a customer or a regulator.
Agents can sometimes work through interfaces designed for humans, but it’s fragile and it breaks when the interface changes. We’d rather tell you it isn’t practical than build something that fails silently in six weeks.
Where two reasonable people would disagree, an agent will pick one and sound certain. That’s worse than no answer.
Agents are genuinely useful for a narrower set of tasks than the marketing suggests, and properly useful for those. We’d rather build three that work than fifteen that need constant supervision — which defeats the purpose.
Which tasks genuinely need judgement, and which are simple rules in disguise.
Whether the systems involved can actually support an agent reliably.
Mapping where a human confirms, before anything irreversible happens.
Built with defined boundaries, so the agent can’t act outside its remit.
Run alongside the manual process until it’s demonstrably reliable.
Ongoing review of decisions made, with boundaries adjusted as needed.
An honest answer on whether an agent suits the task.
Explicit limits on what the agent can and cannot do.
Confirmation steps before anything irreversible or externally visible.
A record of every decision made and why, reviewable afterwards.
Running alongside the manual process until reliability is proven.
API access and credentials registered to you, never held here.
It probably isn’t right yet if simpler automation would do — that’s cheaper and more reliable, and it’s the most common thing we recommend instead. Also not right where errors are expensive and hard to detect, or where nobody has capacity to review what the agent does.
It’s also the phase we’re most likely to advise you against starting with. Not because it doesn’t work, but because the businesses that get value from agents are usually the ones that already fixed the simpler things — and the ones that haven’t get more from doing that first.
Swipe to compare
| Feature | Most Popular Standard | Advanced | Custom |
|---|---|---|---|
| Setup | £5,500 | £9,500 | From£15,000 |
| Scope | Single agent, defined task, one or two systems | Multiple agents or complex multi-system workflows | Bespoke, including regulated environments |
| Checkpoints | Human confirmation on key actions | Configurable approval layers | As specified |
Feasibility assessment: free — including the answer that an agent isn’t right for this.
Ongoing monitoring and review: £450–£950/month. LLM usage billed at cost, in your name. Agent workloads use more than chatbots — typically £40–£200/month, estimated before you commit.
Simpler, cheaper, and usually the right first step.
The structured data an agent needs to work with.
A chatbot answers a question. An agent completes a task — retrieving information, updating records, deciding what happens next, and moving between systems to finish the job.
Often enough that high-stakes actions need a human checkpoint. The shadow-run period exists precisely so you see the real error rate on your own tasks before relying on it.
Boundaries are defined explicitly and irreversible actions require confirmation. There's also an audit trail of every decision, so you can see what happened and why.
For a narrow set of tasks, yes. For most businesses, simpler automation delivers more for less. The audit will tell you which category you're in, and we'll say if it's the second.
Typically £40–£200 a month in LLM usage, more than a chatbot because agents do more reasoning per task. Billed at cost, estimated before you commit.
You are, which is why checkpoints matter. Any agent acting on customer-facing or financial decisions without human review is a risk we'd advise against taking.
Get a free automation audit and feasibility assessment. We’ll look at the task, tell you whether an agent genuinely suits it, and recommend simpler automation when that’s the better answer — which it often is.
We reply within one working day.
No spam, no pitch deck.
WhatsApp us