The removal test
Unplug the AI for a week. Whatever stops was the AI's job. Whatever keeps running never was, no matter what the deck says.
I use one question to find out whether a company is actually built on AI or only decorated with it. Turn the AI off for a week. Not a demo pause. Not a soft freeze with exceptions for the important stuff. No agents, no assistants, no models sitting inside the tools people work in. Then look at what stops.
That is the whole test. It takes a week because a single day only proves you had a buffer.
Most companies pass, and passing is the bad news
Here is what usually happens. Work slows down. A few people complain that drafting is miserable again. Somebody's weekly report shows up late and shorter. And then, by Friday, the same things shipped that always ship.
That company passed the test. It is also not an AI company. It is a good company with a convenience layer on top: a faster way to write an email, a faster way to name the columns in a spreadsheet, a faster way to get to a first draft nobody would have written by hand anyway. Real value. A different thing entirely from what people mean when they say the AI runs it.
I do not think passing is shameful. I think misnaming it is. The cost of the wrong name is that you make plans for a company you do not have.
What stopping actually means
People argue with this test by quietly redefining the word stopped, so I keep the definition narrow and boring.
Stopped means the work did not happen. Not slower. Not uglier. Not "we covered the ones that mattered." If ten things shipped in a normal week and two shipped in the test week, then eight stopped. Count them.
Three things that feel like stopping and are not:
- Everything took longer. That is friction. Friction is a cost, not a dependency.
- Quality dropped. That is a tool loss. Your people still produced the work.
- One person worked the weekend and covered it. That is the answer to the test, stated out loud: a person could absorb it.
The last one is where the test gets cheated, almost always without anyone deciding to cheat. Somebody does the agent's job by hand because they want the week to look fine. Fine is not the goal. Accurate is the goal. You are not being graded, you are being measured.
My company fails on purpose
Take the agents out of my company and almost nothing ships. Campaign builds stop. Content stops. The overnight checks that catch a broken page before I wake up stop. Reports do not get written, because the thing that writes them is gone. The queue backs up on day one, and by day three there is no company, only a person with opinions and a long list.
That is not a fragility I am working to fix. It is the shape I chose.
If removing the workforce does not stop the work, the workforce was decorative.
I picked this on the first day, before there was anything to protect, which is the only easy time to pick it. No team was standing there. No process existed that a person owned. Every procedure got written for a reader that never gets tired and never remembers yesterday unless I wrote yesterday down. So the company's knowledge went into files instead of heads, and the files are the thing that actually compounds.
The failure is loud and that is the point. When the agents break, I know inside a day, because the output is missing. In a company that passes the test, an AI initiative can be quietly dead for a quarter and nothing will tell you.
How to run it honestly
The test is easy to run and easy to fudge. These are the rules I use.
1. Name the work before you unplug
Write down what your company produces in a normal week, in countable units: things shipped, pages published, campaigns built, tickets closed, reports sent, calls booked. If you cannot count your weekly output, stop here. That is your first finding, and it is a bigger one than the test.
2. Pick a normal week
Not the week of a launch. Not a holiday week. If you pick a quiet week to be safe, you will get a flattering answer to a question you asked for a reason.
3. Cut at the source
Turn off keys, accounts, and scheduled agent runs. A policy asking people not to use AI is not a test, it is a survey. People will comply with the letter of it in a browser tab you cannot see.
4. Ban heroics, in writing
Tell everyone the goal is an accurate reading, and that covering for a missing agent by hand is the one thing that ruins it. Say it twice. The instinct to rescue the week is strong and it is coming from your best people.
5. Do not pre-load
No stockpiling drafts on the Friday before. No queueing a month of scheduled posts. Anything generated before the cut and released during the week counts as AI output, so either exclude it or do not do it.
6. Count the same units at the end
Same definitions, same person counting. Then write the number down where you will see it again in six months.
Reading the result
The output of the test is one ratio: the share of a normal week's work that still happened without the AI.
- Near all of it. A normal company with good tools. Nothing wrong with that. Plan like a normal company.
- Most of it, with visible pain. Assisted. The AI is load-bearing in a few specific places, and you now know exactly which ones, which is worth the week by itself.
- Very little of it. The agents are the workforce. The interesting question flips: not what AI could do for you, but what happens when a model provider has a bad day.
Every one of those is a legitimate place to be. Only one of them is the thing people claim to be at conferences.
The cheaper version
A full week is expensive, and for most companies it is not worth the revenue to find out. So run it on one procedure instead. Pick the workflow you most believe is AI-run. Cut it for a week. Same rules: name the units first, no heroics, no pre-loading, count at the end.
You will get a smaller answer of the same shape, usually within two days, and you will get it without an all-hands meeting about it.
What the test does not measure
It does not measure whether the work is any good. A company can fail the removal test completely and still ship things nobody wants. All the test tells you is where the work lives. Whether the work is worth shipping is a separate question, and it is decided by the one thing that never moves to the agents: judgment, taste, direction, and the final gate.
Which is why I run both. The removal test tells me the agents are really carrying the volume. Reading their output tells me the volume is worth carrying. If you want the practical end of this, here is how I would start today with one procedure.
- The era · the prediction nobody has reached yet
- What stays human · the four things I never hand over
- Headcount Zero · the whole idea, on one page
I write field notes when something is worth writing down: what an agent got right, what it broke, what I had to rewrite because of it. Leave an address and they find you. Occasional and unscheduled.
Noted. The next set of field notes goes to that address.
No spam. Used only to send the notes.