The second pilot is the one that tells the truth

First AI pilots succeed because everyone is watching. What the second one is for, why nobody budgets for it, and how to run it without the theater.

There is a meeting I have sat through perhaps a dozen times now. The first AI pilot has wrapped, the deck is glowing, and someone senior asks the only question that matters: "so do we roll it out?" And the room says yes, because the room has spent three months making sure the answer would be yes.

I want to talk about why that yes is usually premature, and why the cheapest insurance in enterprise AI is a second pilot nobody claps for.

First pilots are performances. Not dishonest ones, mostly. But think about who is in them: volunteers, usually the keenest engineers on the friendliest team, with an executive sponsor checking in weekly and a vendor solutions architect on speed dial. Everyone involved knows the initiative is being watched. There is a name for what that does to measurements, and it has been documented since factory lighting studies in the 1920s: the Hawthorne effect. People perform differently when they know they are the experiment.

So the first pilot answers a narrow question: can this work here, under the best conditions we can manufacture? That is worth knowing. It rules out the tools that fail even with a tailwind, and I have watched a few manage exactly that (the stall patterns have not changed much this year). But it says nearly nothing about the rollout, because the rollout will not have volunteers, a sponsor in the room, or the vendor's best engineer answering Slack messages at 9pm.

The second pilot asks the question you are actually funding: does this work with a team that did not ask for it? Pick a group with average enthusiasm, a manager with no stake in the outcome, and the standard support channel (the ticket queue, not the solutions architect). Keep the sponsor away. Measure the same things you measured the first time, against the same baselines you took before any of this started, and then compare the two pilots to each other, not just to the before times.

The gap between pilot one and pilot two is the most useful number in your whole AI program. It is the tax the tool pays when the spotlight moves on. Small gap: the value is probably real, roll it out and spend your energy on training. Big gap: what you had was a performance, and scaling a performance costs a fortune. Fifty seats of enthusiasm do not become five thousand seats of enthusiasm. They become five thousand licenses and a renewal meeting.

Nobody budgets for the second pilot. I understand why. It feels like paying twice for the same answer, the calendar pressure is real, and the first pilot's deck is sitting right there, glowing. But look at the arithmetic. A second pilot costs a quarter of the first one; the templates exist, the integration is built, the measurement plan is written. An enterprise rollout that quietly dies in month eight costs twenty times that, plus something harder to buy back: the organization's willingness to try the next thing. (I have seen that willingness spent. It does not refill quickly.)

One more thing, because it comes up in every one of those meetings. Running a second pilot is not slowing down. The teams I have watched do this well ran it in six weeks, in parallel with contract negotiation, and walked into the renewal conversation knowing exactly what the tool does without a tailwind. The teams that skipped it found out anyway. They just paid enterprise pricing for the lesson, and the lesson arrived with a three-year term attached.

The first pilot tells you what the tool can do. The second one tells you what it will do. Buy the second answer before you buy the seats.