
Evaluating AI vendors when every demo is magnificent
The demo is the vendor's eval, run on the vendor's data. Seven questions that separate products from wrappers, and the clause to read twice.
AI vendor demos have converged on magnificence. The interface is clean, the answers are crisp, and the demo data has been curated the way a show home is furnished. Fine; demos were always theater. The evaluation starts when you make the theater run your script.
Bring your own data, or it is not an evaluation. Twenty real documents, tickets or records from your mess, under NDA, in the room. Products survive contact with your data; wrappers audition politely and ask for a second meeting.
Ask what happens when the model underneath changes. The honest answers involve evals, version pinning and regression testing on their side. The dishonest answer is a blink, because the vendor is three prompts on top of someone else's API and your workflow is one model update from behaving differently.
Read the data clauses like they will be enforced against you. Training rights, retention, subprocessors, region. Then the commercial ones: per-seat pricing that assumes enthusiasm, usage pricing that punishes success. Model the renewal-year bill at realistic adoption, not pilot adoption.
Demand an exit in writing. Your prompts, configurations, fine-tuning data and logs: exportable, in named formats, at contract end. A vendor who resists this question has answered it.
Run the evaluation on your data or do not call it one. The demo corpus is curated the way a show home is furnished. Insist on a two-week trial against your own documents, your own tickets, your own edge cases, with your own staff driving. Vendors who agree are common; vendors who insist on driving the trial themselves are telling you where the magic lives. Bring ten of your ugliest real examples (the scanned PDF, the thread with the forwarded forward, the spreadsheet with the merged cells) and watch what happens. Every product handles the clean case. You are pricing the ugly ones.
Ask about the model dependency out loud. Most AI products are a workflow wrapped around someone else's model, which is fine and worth paying for, but you should know which model, what happens when the provider deprecates it, and whether the vendor's margin survives their upstream price changes. A vendor who cannot discuss this calmly is one upstream announcement away from being a different product. Two questions sort the category: "which model versions did you migrate through last year" (fluent answer good, stammer bad) and "what in the contract happens if your model costs double" (a real answer here is rarer than it should be).
References, but the useful kind. Skip the two happy logos the salesperson offers and ask for a customer who churned, or at minimum one who deployed to more than a thousand seats. The churned reference call is occasionally granted and always illuminating; the refusal is illuminating too. On the call, ask one question above all: "what do you know now that you wish you knew at signature". Nobody answers that with marketing.
And weigh the integration honestly. The build-versus-buy line moves as your team learns: what needed a vendor last year is sometimes four API calls this year, which argues for shorter contracts than procurement prefers. If all of this reads as adversarial, it is the opposite. The vendors worth keeping enjoy informed buyers, because informed buyers deploy successfully, renew for reasons, and become the reference calls that close the next deal. The evaluation ritual above filters for exactly the vendors who will still be pleasant to work with in year two, when the demo team is long gone and the relationship is a support queue and an invoice. Buy from companies that survive your questions. They are buying a customer who survives theirs.
The magnificent demo will still be magnificent next quarter. Your leverage will not survive a three-year term signed during the honeymoon.