Abstract artwork of three warm bands progressing from sparse to structured

The first 90 days of deliberate AI adoption

A quarter is enough to move from enthusiasm to evidence, if the sequence is right. The month-by-month plan we hand clients after the audit.

The 90-day plan at the end of our readiness audit is always built from the same skeleton, because the sequence, unlike the contents, barely varies. Here it is, for teams who want to run it themselves.

Days 1 to 30: make the present legible. Publish the one-page usage policy with its amnesty clause, and run the shadow inventory it makes possible. Take the productivity baseline while nothing is changing yet. Pick the two pilots for next month using the audit's boring criteria: a named workflow owner, data already in usable shape, success as a number. Resist the urge to pick three. The third pilot is where focus goes to die.

Days 31 to 60: run small, measure honestly. Both pilots live, each with a weekly review that looks at the metric rather than the vibes. Sanctioned tools reach the shadow users first (they were your requirements document; now they are your power users). Role-shaped training starts with the pilot teams only. Nothing rolls out company-wide this month, however good the demo feels, because company-wide is where unmeasured things become unkillable.

Days 61 to 90: decide in public. Each pilot ends in one of three verdicts, announced with its numbers: scale it, fix it and rerun, or kill it with thanks. The kill, done openly, buys more credibility than either success; it proves the numbers mean something. Then the next two pilots enter the pipeline, and the cycle repeats until it stops being a program and starts being how the company works.

The meetings that hold it together. The whole apparatus runs on two recurring calendar entries, and resisting the urge to add a third is half the discipline. A weekly thirty-minute pilot review: the metric, the blockers, nothing else, and it ends early most weeks. A monthly hour with whoever owns the budget: what the numbers say, what dies next, what needs a decision only they can make. No steering committee, no working group, no program office. I have watched a twelve-person committee spend a quarter designing the governance for pilots that two people could have simply run in the same quarter, and the committee's slides were, to be fair, excellent.

What goes wrong, specifically. Three failure patterns account for most broken first quarters. The pilot owner is reassigned in week five, because the plan assumed calendars freeze; name a deputy on day one. The baseline gets skipped because the pilot is "obviously valuable", and month three arrives with a great story and no number, which the CFO files under stories. Or legal discovers the rollout in week seven and everything stops; the fix costs one coffee in week zero, showing them the policy and asking what would make them comfortable. Legal teams block what surprises them, not what consults them.

Why two pilots and not one. A single pilot becomes a referendum on AI itself: its failure kills the program and its success gets credited to everything except the workflow fit. Two pilots make comparison possible, and comparison is where the learning lives. When one scales and one dies, the difference between them, usually data readiness or owner attention, becomes the most valuable sentence in the quarter's report, and it points the next pair of pilots somewhere smarter. One is an anecdote. Two is the beginning of a method.

What you should be able to say at day 91. Something with this shape: "Two pilots ran. One cut claim-drafting time 30 percent and scales next quarter with these guardrails. One missed its number, here is why, and here is what we changed for the rerun. The policy has been live for a quarter, four questions came in, none required a lawyer." Notice what that paragraph contains: verbs, numbers and dates. Notice what it does not contain: the word "transformation". Boards fund the first kind of paragraph indefinitely.

If you inherit a program already in motion. Half the teams who pick up this skeleton are not starting clean; they arrive with four half-running pilots, a paused policy draft and a vendor trial someone signed at a conference. The salvage sequence is the same skeleton run backward: freeze new starts, apply the three-verdict test to everything already running (most inherited pilots take the "fix and rerun" door once they finally get a metric), publish the policy this month regardless of its imperfections, and take the baseline now, labeled honestly as mid-stream. A mid-stream baseline is worth less than a clean one and infinitely more than the alternative, which is arguing about attribution forever.

A quarter from now you will have evidence, either kind. That beats what most AI programs have at the one-year mark, which is momentum and a feeling.