Ai agents: what they can and cannot do yet.

Everyone reading this has been pitched one by now. The pitch is broadly the same wherever it comes from: describe what you want, and software goes away and does it, using the same tools your staff use.

We use these tools every day on real work, and what follows is where they stand, put as honestly as we can manage.

What an agent is, and what it is not

The word covers a lot of ground, but used carefully, an agent is a language model that has been given tools and permission to use them in a loop. It reads a request, decides on a step, takes the step, looks at the result, and decides again. The loop is what makes it an agent rather than a chatbot: it can take actions, see the consequences, and change course.

A great deal being sold as agents is not this. Some of it is a scripted sequence with a language model producing text at one point in the chain, which is a useful thing but an ordinary one, and priced differently. Ask which you are being shown, because it changes both what you can expect and what can go wrong.

What they are genuinely good at

They are good at turning one form of information into another, taking something unstructured and making it structured, or the reverse. Twenty rambling enquiry emails become a table with the right fields filled in, and a set of meeting notes becomes a list of commitments with names against them. This is the strongest thing they do and it is consistently underrated, because it sounds mundane.

Then there are first drafts of anything formulaic, valuable not because the output is good but because a competent starting point now costs almost nothing, and editing is faster than beginning. That is real value even when the draft needs heavy work.

They read more than a person will. If you have four hundred support tickets and a question about them, an agent reads all four hundred where a person reads thirty and generalises. For finding themes and outliers in a body of text nobody has time for, they are already better than the realistic alternative, which is usually nobody looking at all.

And they grind through long, tedious, well-defined sequences. Where the rules are clear, the tools are available and the work is only tedious rather than difficult, an agent will keep going without getting bored, and boredom is where people make their mistakes.

Where they reliably fall over

They do not know what they do not know. This is the central limitation and most of the others follow from it. A model produces the most plausible continuation, and it has no separate faculty checking whether that continuation is true. It cannot reliably tell the difference between recalling something and constructing something that sounds like the thing it would recall.

They lose the thread on long tasks, so give an agent a fifteen-step job and quality falls off across the sequence. A small misreading at step three does not get corrected at step nine, it gets built on, and the practical limit is not how clever the model is but the length of the chain.

They are inconsistent, so running the same task twice may get you two different answers, both reasonable. For drafting that is fine. For anything where the same input has to produce the same output it is disqualifying, and it is why a great deal of finance and compliance work stays with ordinary deterministic software.

They are poor at knowing when to stop and ask. A person doing an unfamiliar task hits a point of real uncertainty and comes to find you. An agent, unless carefully constrained, proceeds. It makes a reasonable assumption and carries on, and the assumption does not arrive labelled as one.

And they have no stake in the outcome. That sounds like a philosophical point and it is a practical one. A bookkeeper who cannot reconcile an account gets uncomfortable and escalates, because being wrong has consequences for them, and nothing in the software feels that.

This is the part the pitch skips, and it is the part that decides what you can safely hand over.

What the failure actually looks like

This is the part the pitch skips, and it is the part that decides what you can safely hand over.

The failure is not an error message, a crash, a blank screen or a job that stops halfway with a red warning, all of which are the easy case because you know. The characteristic failure of these systems is a complete, confident, well-formatted, plausible wrong answer, produced in the same tone and the same layout as the correct ones and delivered at the same speed.

The invented figure sits in a column of real figures. The summary of a document describes what such documents usually say rather than what this one said. The cited policy sounds exactly like your policy and is not in your handbook. In every case the output passes the glance that most work actually receives, and passing the glance is the problem.

The consequence is that the value of an agent is capped by your ability to check its work. That is the single most useful thing to hold on to. If verifying the output takes as long as doing the task, you have gained nothing. If verifying is fast and doing is slow, you have gained a great deal. That ratio, rather than the sophistication of the model, is what decides whether a given task is worth handing over.

What has to be true before you hand a task over

There are four conditions, and they are demanding on purpose.

The output is checkable quickly by somebody who would know. Either a person verifies it in a fraction of the time it took to produce, or something deterministic checks it: a total that has to reconcile, a field that has to match a record you already hold.

Being wrong occasionally has to be survivable, not tolerable in theory but survivable in practice, at a realistic error rate, when the wrong answer is confident and looks right.

The task has a boundary, so the agent reads what it needs and writes only where you have decided it may. Anything that sends externally, moves money or deletes should sit behind a person. Not because the technology is untrustworthy in principle, but because the failure mode is a confident wrong action, and confident wrong actions are much worse when they cannot be undone.

And somebody owns the output, not the tool and not the supplier, but a named person whose work it is once it leaves the building. The moment nobody owns it, nobody checks it, and the whole arrangement rests on checking.

Where that leaves it

What you can buy today is a fast, tireless, capable and unreliable assistant that will never tell you it is unsure. That combination is genuinely valuable, and it is valuable in a specific shape: give it work that is laborious to do and quick to check, and keep a person on anything expensive to get wrong.

The businesses getting real value from this are not the ones who moved fastest. They are the ones who worked out what they could check, handed over exactly that, and left the rest alone.

The next step is a list rather than a project. Write down the tasks that are laborious to do and quick to check, and most businesses find three or four. Start with the one where a wrong answer costs least, and put the checking step in writing before you build anything.

That is the shape of most of what we build with these tools. Using them daily is mostly how we know where they break, and the work we take on is the work that survives the checking test rather than the work that demonstrates well.

Azeem Hadi, founder of Baseops

Azeem Hadi

Founder @ Baseops. I've been crafting websites, creating brands and building systems for nearly 20 years.

Got a project in mind?

Or just want to talk through an idea? We are always happy to chat, no strings attached.

Thirty minutes, direct with the founder. No pitch deck.