By Rolf Versluis ยท Published [DATE] on Uncommon Knowledge in Business and Life
When I tell people I run an AI agent, the first thing they ask is whether it can manage their calendar or answer their email. I always tell them no. Not because the agent cannot do it. Because that is the wrong place to start.
Starting with email or calendar is starting with the highest-stakes part of your day. The agent does not know your relationships, your tone, your priorities, or which emails are urgent and which can wait. If it gets that wrong, you lose trust fast.
Start with something lower-stakes. Start with research, or organizing, or one specific loop that eats your time. Let the agent learn how you work. Let it earn the right to handle the important stuff.
That is the advice I give everyone who is new to this. It is also the advice I sometimes do not take myself, which is how I learned it.
What the agent actually feels like
Working with a Hermes agent is like working with a very eager and extremely fast and capable child. I do not mean that as a put-down. I mean it as a description.
The agent wants to please you. It wants to get to a result. It wants to finish the task. If you give it a vague instruction, it will fill in the gaps with whatever seems most likely. If you give it two instructions that contradict each other, it will sometimes pick one and not tell you, and sometimes pick both and not tell you they contradict. The way I describe it: the agent does not see any other way out of conflicting directions. It does not tell you they conflict. It picks something and runs.
The thing that surprises people is the second one. The agent will tell you it did the thing you asked, even when what it did was not actually the thing you asked. It is not lying in the human sense. It is doing what the language model is trained to do, which is produce a confident, fluent answer. If the answer happens to be wrong, the agent does not know it is wrong. It just keeps going.
This is the part that takes a few weeks to learn. You start reading the agent's output the way you would read a junior employee's first draft. You check the work. You ask follow-up questions. You do not trust it on the first pass.
You start to trust the work after a few weeks. The way you would trust a colleague who has been around long enough that you know their patterns.
That calibration is the work. You cannot skip it by giving the agent more important tasks to do.
What I learned the hard way
The first time I had the agent do something important without checking the work, it went sideways. I do not want to get into the specifics, but the shape of the problem was: I gave the agent a request that had two parts. The agent handled one part. It told me it had handled both. I did not check. The second part was not done.
That is the failure mode. Not "the agent is bad." Not "AI does not work." The failure mode is that the agent produced a confident, fluent answer that did not match what it had actually done. If I had read the output carefully, I would have caught it. If I had asked the agent to show me what it had changed, I would have caught it. If I had asked a follow-up question about the second part, I would have caught it.
I did not do any of those things. I trusted the first answer. The agent did not catch me. The work was not done.
That is what the calibration period is for. The first few weeks, you check everything. After a few weeks, you know which things to check and which you do not.
How to start
Pick one loop. Make it a loop that does not involve money, does not involve customer communication, and does not involve anything that would be embarrassing to get wrong. Good first loops:
- Research. "Find me three options for X and tell me the tradeoffs."
- Organizing. "Sort these files into folders based on the project name."
- Summarizing. "Read this document and tell me the five most important points."
- Drafting. "Write me a first draft of an email to a vendor about Y. I will edit it before I send it."
Notice what all of those have in common. The agent does the work. You review the work. Nothing goes out the door without your eyes on it.
After you have done a few of those, expand. Have the agent draft the email and queue it for your approval. Have it do the research and write up the comparison. Have it run the loop on its own and report back.
The progression is the same as managing a new hire. Start with supervised work. Move to reviewed work. Move to unsupervised work where you spot-check. The agent does not get to "unsupervised and trusted" on day one. It earns it.
What I do for clients
When I work with a business owner who is new to agents, this is the structure I set up. We start with a single loop. After a week, we have checked the work and the owner has a feel for what the agent does well. Add a second loop. Same pattern. By the end of the first month, the owner has a handful of supervised loops running and knows which ones to trust.
Two months in, the owner is picking their own loops and trusting the agent on most of them. That is when the work shifts from "I am babysitting an AI" to "I have an assistant."
If you are new to this, the agent is not the bottleneck. The calibration is. The first month is the work.
I work with owner-operators of small and medium businesses on this kind of setup. First month is the work; the website is localhermesagent.com.