AI purchases have a way of being judged by vibes: the demo was impressive, a few people say they like it, the subscription renews. Vibes are how you end up paying $30 a seat for a tool eleven people opened twice. The alternative is small: a baseline, three numbers, and a 90-day verdict. Hours saved or it didn't happen.
Baseline first, or you're guessing forever
Before the pilot starts, measure the task as it works today: how long one instance takes, how many instances a week, who does them. Real numbers from the people doing the work, not estimates from the meeting about the work. "Proposals take about three hours and we do five a week" is a baseline. Without it, every later claim is unfalsifiable, which is precisely how bad tools survive.
The three numbers that matter
- Net time per task. New task time including AI time, prompt fiddling, and, critically, review and correction time. AI that drafts in 30 seconds but needs 40 minutes of fixing saved you 20 minutes, not three hours. Count honestly.
- Correction rate. How often output needed meaningful fixes, and whether anything wrong nearly shipped. This number tends to improve for the first month as people learn the tool, then plateau. The plateau is the truth.
- Adoption. Of the people with licenses, how many used it this week? Tools that work get used without nagging. A licensed-but-idle seat is the clearest signal there is, and it's a number your admin console already has.
The arithmetic, worked once
Five proposals a week, three hours each, baseline. With AI: 1.5 hours including review. Savings: 7.5 hours a week. At $50 loaded hourly cost, that's $375 a week, call it $1,600 a month, against maybe $90 a month of licenses for the three people involved. Verdict: obvious, expand it. Same math on the support-email pilot: 10 minutes saved a day across two people, $35 a month of value against $60 of licenses. Verdict: kill it, or find the version of the task where the model earns more. The point isn't precision to the dollar. It's that a five-minute calculation beats any opinion in the renewal meeting, including ours about where AI helps.
The 90-day verdict
Every pilot gets a decision date on the calendar at birth: expand, adjust, or kill. Expand means the numbers work; roll it to more people or graduate the workflow to a deeper tier. Adjust means promise but friction: often fixable with training, a better prompt template, or cleaner source documents. Kill means the numbers said no, and killing on schedule is a success of the process, not a failure of the pilot. The 90-day habit is what keeps the AI budget pointed at the five jobs that pay instead of accreting subscriptions the way unused licenses always accrete.
Report it like anything else
Hours saved monthly, cost, correction incidents, adoption. Four lines on the same page as your other numbers worth watching. When AI earns a line on the real dashboard, it stops being a novelty and starts being what it should have been from the start: a tool with a job and a scorecard.
Want this handled instead of homeworked? That's the job.
Email us →