I don’t want more automation. I want less unnecessary human work.
As AI agent workflows become more autonomous, I’m becoming more interested in which human involvement reflects real judgment and which is only operational plumbing.

For a while, one of the easiest ways to think about progress with AI agents was to ask how much more they could do without me. Could they take a task further? Could a workflow run longer before it needed intervention? Could more of the execution, checking, handoffs and recovery happen without a person pushing it along?
Those are still useful questions. But recently I have noticed that they can point at the wrong target. A workflow can become more automated and still demand a surprising amount of human attention. The human may no longer be doing the main task, but is still checking whether it finished, carrying information from one place to another, retrying something ordinary, relaying a result, or repeatedly telling an agent to continue. The work has changed, but the person is still helping to keep the machinery moving.
In the previous Journal entry, I wrote about a different problem with very smooth AI collaboration: some friction is waste, but some resistance contains useful information. Removing every source of disagreement can make it easier to keep moving in the wrong direction. That left me with a more practical question. If some human involvement really is valuable, which parts deserve to remain?
I have started asking myself a much simpler question when I look at an AI agent workflow:
Why does a human still need to be involved here?
Sometimes I can answer that immediately. I may need to choose between meaningful product trade-offs, challenge the direction, decide whether the available evidence is sufficient, accept a real risk, resolve genuine ambiguity, or make a business or customer decision. There are also actions where authorization itself matters because somebody has to take responsibility for what happens next. In those moments, I am not interested in removing the human just to increase an autonomy score. There is an actual decision being made.
Other interventions are much harder to defend. Checking again to see whether a task has finished is not product judgment. Copying state between systems is not oversight. Retrying an ordinary recoverable failure, relaying information between tools, coordinating a routine handoff or telling an agent to continue usually does not become more valuable because a human performed it.
A lot of this work feels less like human oversight and more like human plumbing. That distinction has made me more skeptical of the phrase “human-in-the-loop.”
The phrase can be useful, but by itself it tells me only that a person appears somewhere in the workflow. It says very little about what that person is actually contributing. An approval step can represent an important decision. It can also mean that somebody has to press a button before the system is allowed to continue.
If the person has no meaningful judgment to exercise, has no useful alternative to consider, and is expected to approve the same ordinary case again and again, the workflow technically contains a human without necessarily gaining much human oversight. That can even create a misleading sense of safety. We can point to the approval and say that a person was involved, while the person may actually be functioning as a manual continuation mechanism.
I have noticed a similar problem with monitoring. There are situations where I genuinely need to know that something has changed, failed, become ambiguous or reached a consequential decision. There is much less value in repeatedly looking at routine progress simply because nothing else will tell me whether I need to act.
The difference is important because human attention is not free just because the individual action is small. A few seconds to check something, another minute to move information somewhere else, another interruption to confirm that ordinary work can continue — none of these looks particularly expensive in isolation. But when a workflow becomes longer and more agentic, those small touches accumulate. The person is no longer doing the execution, yet still cannot really leave it alone.
That does not feel like the leverage I was trying to create. At the same time, noticing all of this creates an obvious temptation: remove more human intervention. I do not think that is the answer either.
If a real ambiguity exists, removing the person does not make the ambiguity disappear. Something still resolves it. The decision may simply move into a model response, a default behavior or an assumption made earlier in the workflow. The same applies to risk and direction. Automation can remove work without removing responsibility. A workflow can execute a bad product decision extremely efficiently, and extra autonomy does not make the original decision better.
This is where my thinking has shifted. I am becoming less interested in how autonomous a workflow can become as an objective by itself. I care more about whether each remaining human touch has a reason to exist.
If I am there because the workflow needs somebody to monitor routine progress, copy information, perform ordinary retries or keep handing work from one step to another, I increasingly see that as something worth designing away where it can be done safely. If I am there because a genuine decision has arrived — something involving judgment, ambiguity, risk, authority, disagreement, a customer or a business choice — then removing me may make the system smoother without making it better.
The boundary is not always obvious, and I do not think it stays fixed. Something that requires judgment today may become sufficiently understood and repeatable that it can be handled routinely later. An operation that is usually routine can also encounter an unusual case where a person suddenly does need to become involved. The useful question is not simply whether a particular step is “automated” or “manual,” but what kind of work is actually happening at that moment.
One effect of reducing routine human touches has surprised me: the remaining human touches become more meaningful. When I am not being interrupted to babysit ordinary execution, an interruption can carry information. It can mean that there really is a choice to make, uncertainty to resolve, risk to accept or direction to challenge.
That is a much better use of human attention than asking somebody to approve everything just because having an approval step sounds responsible. It also changes what I consider a good AI agent workflow.
I do not need the system to prove its sophistication by never asking for help. A workflow that never surfaces uncertainty or consequential decisions can be just as questionable as one that asks for permission after every routine action.
What I increasingly want is for the ordinary continuation to become ordinary. Monitoring, retries, routine coordination and mechanical handoffs should require less attention as the workflow becomes capable of handling them reliably. The moments that genuinely depend on human judgment should become easier to notice precisely because they are no longer buried among dozens of meaningless interventions.
I still do not have a clean formula for where that line belongs. I am not sure there should be one that applies to every product, every agent or every workflow. But I have a question that is becoming much more useful than “How autonomous can this become?”
When the workflow asks for a human, can I explain why that particular moment deserves human attention?
If I can, there is probably something meaningful for the person to do. If I cannot, I increasingly treat that as a workflow problem rather than automatically calling it oversight.
The goal is not to make humans disappear. I want it to become increasingly obvious why a human is there when the workflow asks for one.