What stays human
Four things stay with me: judgment, taste, direction, and the final gate. Everything that is volume goes to the agents. I have tried handing over the four and it goes badly every time.
The split I run on is simple enough to say in a sentence. The agents carry volume. I make the calls.
Volume is everything that scales by doing more of it: drafting, building, checking, formatting, gathering, reporting, the ninth version of a thing that already exists in eight versions. That work used to consume a company's whole payroll, and it is the work an agent workforce is unreasonably good at.
The calls are different in kind, not in size. They do not get better by being done more times per hour. Four of them stay with the person, and I want to be precise about what each one is, because the vague version of this argument is the one people use to feel safe.
Judgment
Judgment is deciding what is true enough to act on.
Every day something arrives that looks like a fact and is not one. A number that came from a broken query. A summary that is accurate about the wrong week. A confident recommendation built on a source that has been dead since March. The output is fluent, internally consistent, and wrong, and fluency is exactly what makes it dangerous.
What I bring is not more intelligence than the machine. It is a different kind of suspicion. I know what a report looks like when the pipeline broke. I know which numbers move together and I get uneasy when one of them moves alone. That suspicion is built from having been burned, and being burned is a human input that does not transfer through a prompt.
So I read the work. Not all of it forever, but enough of it that my nose stays calibrated. The moment I stop reading, the calibration goes, and then I am an approval button with an opinion.
Taste
Taste is picking between two acceptable versions.
Agents are excellent at producing things that are fine. Fine is a real achievement and it is not what wins. Given twenty finished options that all meet the brief, someone has to know which one is actually good, and that call is not derivable from the brief, because if it were derivable I would have written it in the brief.
I can specify what is required. I cannot specify what is good. So far I can only demonstrate it, one correction at a time.
This is the slowest thing to transmit and the most valuable thing I own. Every time I reject a version and say why in a sentence that ends up in a document, a little of my taste moves into the system. Over months the average output rises, and the corrections get more specific and less frequent. It is genuinely a form of teaching, done in writing, to a reader that will never remember it unless I write it down.
It never finishes. Taste is partly about what the market wants this month, and the market keeps moving.
Direction
Direction is choosing the frame.
Agents optimize inside a frame beautifully. Give them a goal and rails and they will push against the rails all day. What they do not do is stop and ask whether the goal was worth having. Point a capable system at the wrong objective and it gets you to the wrong place faster than a slow team ever could, and it will look like progress the entire way.
Which market, which product, which customer, what we refuse to do, what we are willing to be bad at. Those are mine. They come from things the system does not have access to: what I am willing to live with, what I find interesting enough to still be doing in three years, what I think is coming.
Direction is also where the leverage is highest and the feedback is slowest. A bad procedure shows up in a day. A bad direction shows up in a quarter, sometimes later, and by then a lot of excellent work has been done in the wrong direction.
The final gate
Nothing goes out that a person has not approved.
That is the rule, and keeping it honest at volume is harder than stating it. Reading every word of everything stops being possible quickly, and pretending otherwise turns the gate into theater. So the gate has layers.
- Hard rails in code. The things that must never happen are enforced by the system, not by my attention. Attention fails at 2am. A check does not.
- Full review where it is irreversible. Anything public, expensive, or hard to take back gets read completely, every time, no exceptions for being busy.
- Sampling where it is reversible. For high volume, low stakes work, I read a slice, and I read a bigger slice when something feels off.
- A named person on the outcome. Not the agent. Me. If it goes out wrong, that is my mistake, and I want the incentive structure to make that obvious.
The gate is the difference between an agent workforce and a machine that publishes. If you would not put your name on the output without reading it, you do not have a workforce, you have a liability with a schedule.
Why removing yourself entirely is not the point
People assume the end state of this is a company with no human in it at all. I do not want that, and I do not think it is available.
Not available, because someone has to hold the four. Remove the judge and the system optimizes into nonsense with total confidence. Remove taste and the output converges on the average of everything, which is the definition of forgettable. Remove direction and the frame drifts to whatever the last instruction implied. Remove the gate and the first bad day becomes a public one.
Not wanted, because the point was never absence. The point was to stop spending myself on volume so I have room for the calls. I did not build this to work less. I built it so the hours I do work go into the part that only a person can do, and so the company keeps producing on the days I am not at the desk.
There is a version of this that fails badly, and it fails by automating the judgment too. It looks efficient right up until the moment nobody in the company can tell whether the output is any good, and by then there is nobody left who ever could.
Making judgment cheaper instead of scarcer
The real constraint in a company like mine is not how much the agents can produce. It is how much I can approve. As output grows, I become the bottleneck, and no amount of extra capacity fixes a bottleneck downstream of it.
So the work is not to make more. It is to make each call cost less.
Three things move that needle. Write the standard down, so a judgment I already made does not need making again. Turn any correction I give twice into a rule the system enforces itself. And build checks that surface the exceptions, so my attention lands on the things that are unusual rather than on a wall of things that are fine.
That is the whole loop, and it never runs out of work, because every rule I add reveals the next thing I was deciding by instinct without noticing.
If you want the practical starting point, it is one documented procedure, verified every run. If you want to know whether the agents are really carrying the volume in the first place, run the removal test.
- Start · the first procedure and how to verify it
- The era · a prediction nobody has reached yet
- Headcount Zero · the whole idea, on one page
I write field notes when there is something worth reporting: what an agent got right, what it broke, what I had to rewrite because of it. Leave an address and they find you. Occasional and unscheduled.
Noted. The next set of field notes goes to that address.
No spam. Used only to send the notes.