- delegation
- productivity
- operations
Delegate to an AI agent like you would to a new hire
Nobody hands over the keys on day one. A four-level model for handing tasks to an AI agent without losing control — or your nerve.
You don't hand a new hire the keys on day one. You show them the ropes, let them try, check their work, correct it, and gradually let go. Delegating to an AI agent works exactly the same way — but almost nobody treats it that way. Either it's expected to be flawless from the first message, or it's left alone far too early. Here's a simple way to delegate in stages, without losing control or holding your breath.
The mistake: treating it like software
Software gets installed and works. An agent doesn't: an agent makes decisions, and decisions have to be learned. Treat it like just another app and you'll expect it to be right on the first try — then the moment it sends a customer something odd, you switch it off and go back to doing the work yourself. Understandable reaction. It's also the fastest way to bin a tool that would have taken hours off your plate within a fortnight.
Flip it around. If a new employee gets the collection times wrong on day two, you don't fire them; you tell them how it works and move on. What you do do is not let them sign off quotes yet. That distinction — what they can do alone, what goes through you — is exactly what an agent needs. And it's worth writing down.
Four levels of delegation
Instead of the binary question ("do I automate this or not?"), use four levels. Every task sits at one of them, and moves up or down based on how it behaves.
Level 0 — Watch. The agent touches nothing. It reads incoming messages and suggests a category or a summary. You still reply. It's a zero-risk way to see whether it actually understands your business. A week is usually enough.
Level 1 — It drafts, you send. The agent writes the reply and leaves it ready. You read it, tweak it if needed, hit send. This is where the real learning happens: every edit of yours is a new rule. Go five days without editing anything and that task is asking to be promoted.
Level 2 — It acts and flags. The agent replies on its own but logs everything and pings you when something falls outside the script — an angry customer, a question its playbook doesn't cover, a request outside opening hours. You're not approving message by message; you're supervising the pattern.
Level 3 — It just acts. Reserved for repetitive, well-bounded, low-risk tasks: confirming an appointment, sending the address, answering "are you open on Saturday?". Nobody needs to review that.
The goal isn't to get everything to level 3. The goal is for each task to sit at the level it deserves, and for you to know which. An availability question can live at level 3 while a complaint about a broken order stays at level 1 forever. That's a healthy setup, not a failure.
What should never be promoted
Some decisions shouldn't be delegated even if the agent gets them right every single time — because the issue isn't accuracy, it's who answers for the outcome.
- Money outside your price list. Applying your published rates, fine. Negotiating a discount or closing a big quote, no.
- Commitments you can't undo. Confirming a date that blocks half your team, accepting a late return, promising a rush delivery.
- Conflict. Angry customer, public complaint, your own mistake. A correct-but-cold reply does more damage here than a plain "I'll call you in ten minutes."
- Anything with legal or medical consequences. A diagnosis, tax advice, a contract interpretation. The agent can gather the details and set the table; you give the answer.
A rule that holds up: if getting it wrong would cost you an apology, delegate it. If it would cost you a customer or an invoice, approve it yourself.
And one honest caveat about this whole article: if a task happens three times a month, don't automate it. Delegating properly costs time up front — writing the rules, reviewing that first week — and you only earn that time back through volume. Three messages a month? Just answer them.
Delegate the judgement, not just the task
The most common failure isn't technical. It's asking an agent to "handle customer messages" without ever explaining what a good reply looks like in your business.
A new hire picks up, in their first week, a pile of things nobody has ever written down: that maintenance clients get answered first, that you never confirm anything on a Saturday because the warehouse is shut, that if someone asks about the old model you have to warn them there are no spare parts left. None of that lives in a manual. It lives in your head.
An agent won't guess it. So write it down — and no, you don't need a forty-page document. This is enough:
- The ten questions you get most, with the good answer. Not the generic one: yours, with your lead times and your conditions.
- Three or four priority rules: who gets answered first, and why.
- The hard limits: what you never promise, never discount, never confirm without you.
- The escalation line: exactly what it says when something is beyond it. "Marta is looking at this and will get back to you before 6pm" reassures people far more than a smooth non-answer.
Half an hour of your time. It's the difference between an agent that helps and one you're forever cleaning up after.
How to know it's working
Promote tasks on evidence, not on impatience — or on nerves.
For the first week, track two numbers per task: how many replies you had to tweak, and how many you had to rewrite from scratch. The second one is the one that matters. A style edit is noise; a reply that stated something untrue is a missing rule.
A workable threshold: five consecutive days with no substantive corrections on a task, and that task moves up a level. And in reverse — when something serious slips through, demote that task, and only that task, until the rule is clear. Don't shut the whole thing down over one bad corner. You wouldn't sack someone for fumbling a single order.
After a month, the question stops being "is it accurate?" It becomes a more useful one: how many hours have you stopped spending on this? If you can't answer that, the automation isn't really working — however accurate it is.
Start with one thing
Don't delegate "customer service". Delegate confirming tomorrow's appointments, at level 1, this week. When it goes five days without you touching it, move it to level 2 and pick up the next task. In a month you'll have four or five tasks spread across the levels and a very clear sense of where your trust ceiling sits.
That model — you approve, they execute — is how we build agents at Yaqbot: they start by proposing and earn autonomy one task at a time, never all at once. If you want to see what it looks like day to day in a small business, the operations agent is the closest fit: it gathers, prepares, flags, and leaves the decisions where they belong.
