Why our agents ask permission before the actions that matter
This year's story is autonomy. Agents chaining tasks together, opening their own browser, writing into your systems and closing the loop without anyone looking over their shoulder. The technology really is arriving.
The question still to be answered is what happens on day forty. And there, in the projects we've built ourselves, what separates a striking pilot from something still running months later has nothing to do with how powerful the model is, but with where you decided to put the handbrake.
The problem isn't that it gets things wrong, it's how fast
A person who enters a figure wrongly makes one mistake and, with a bit of luck, somebody catches it that same afternoon. An agent that enters it wrongly makes the same mistake four hundred times before lunch, without tiring and without setting off any alarm, because from inside the system everything looks fine: tasks complete, the queue goes down and the logs say it's done.
That's why the useful question when you're designing an automation isn't whether the agent can do something on its own, since it almost always can. The question that actually matters is what would happen if it did it wrongly five hundred times in a row. If the answer is that you fix it in ten minutes, go ahead; if the answer is that you're calling two hundred customers to apologise, that's where a human approval belongs.
Where we draw the line
It isn't an abstract principle that shifts from project to project, but a list we apply the same way every time and put in writing before we build anything.
- Alone, without asking: reading, classifying, summarising, drafting or moving information between internal systems. Everything that's reversible and doesn't leave the building.
- With human approval: anything that goes out with your name on it, touches money, deletes something or writes into a system other people depend on. An email to a customer belongs here, and so does an invoice.
- Never, even if it technically can: deciding about a person, whether that's screening a CV or scoring somebody's creditworthiness. The EU regulation classes it as high risk.
Asking permission doesn't mean approving things all day
That's the usual misunderstanding and it's a fair worry, because if every step needs signing off then you haven't automated anything, you've just swapped one job for another. The trick is that the agent arrives with the work done right up to the last click: it drafts the whole email, prepares the whole invoice and leaves the reply ready, so all it asks the person for is a yes or a no with the consequence in plain sight.
In practice that turns an hour of work into two minutes of review, which is where nearly all of the saving is, and in exchange it rules out the one scenario that genuinely hurts, which is something going out with your name on it that nobody looked at first.
What this has to do with this week
Quite a lot, which is why we're writing it now. Article 50 of the EU regulation becomes applicable these days, requiring you to say when there's a machine behind something. It's the same idea seen from the legal side, because however autonomous the system is there's always somebody who answers for what it does, and that somebody has a name.
An agent that asks permission on the sensitive things isn't a less capable agent, it's one where you know exactly where the signature is.
What we'd rather not promise you
We aren't going to tell you full autonomy won't arrive, because it probably will and sooner than we think. What we're saying is that today, with what exists, building an automation with no control points amounts to betting that nothing odd will happen for months, and when it does, the cost ends up being paid by the client.
So we'd rather start with the boring version, handbrake where it belongs, and take it off later once we've watched the system run for a few weeks. It looks less impressive in a deck and considerably better in March.
Got a process in mind and no idea where that handbrake would go? We'll look at it together in the free audit.