Atlas / GOVERN & LEAD / Strategy / Vendor Selection
DEEP-DIVE · STRATEGY

Selecting AI Vendors: A Procurement Playbook

Deciding to buy narrows nothing until you choose whom. Vendor selection is its own procurement discipline: anchor on your requirements, prove capability on your own data, and win the data and exit terms before the features.

TL;DR
  • Choosing to buy is not choosing a vendor. Selection is a distinct procurement discipline that sits after the build-versus-buy call and after the category map, and it is where most of the durable risk is actually decided.
  • Anchor the evaluation on your own requirements and differentiating criteria before the demos frame your thinking, then prove capability on your own data through a POC or bake-off. That step is the single most important and most-skipped move in the whole process.
  • For regulated buyers the data terms and the exit terms often matter more than the feature list. Scrutinize training rights, residency, retention and no-train guarantees; price the lock-in; keep a gateway as an abstraction seam; and re-run the evaluation as a young, churny market moves.

Choosing whom, not whether

Two decisions usually precede this one, and it is worth being clear that neither of them is this decision. The first is build versus buy versus orchestrate: whether the capability is core enough to build, commodity enough to buy, or best assembled from parts (the case is made in the build-buy-orchestrate deep-dive). The second is the category map: understanding what kinds of platform exist and which class of product the need belongs to (see the platform landscape deep-dive). Both are strategy. Neither tells you whom to sign.

Vendor selection is a procurement discipline, and treating it as a continuation of the strategy conversation is where programs quietly lose money. Once "buy" is settled and the category is named, you are choosing a specific company to take a dependency on: their roadmap, their data practices, their solvency, their support desk at 2am. That is a different exercise, with different failure modes, and it rewards procurement rigor rather than architectural taste.

The stakes are higher here than in a conventional software purchase, for one structural reason. An AI vendor does not just run code against your inputs; it processes your data through models whose behavior you cannot fully inspect, often trains or improves on what it sees unless you stop it, and sits close enough to your knowledge and your decisions that switching later is genuinely painful. The feature comparison that dominates most selection decks is the least durable part of the choice. The parts that outlast the demo, the data terms, the exit path, the vendor's ability to still exist in three years, are the parts a hype-driven process skips.

This article is deliberately advisory rather than a directory. It will not name products or rank suppliers, because any such list is stale within a quarter. It offers instead a repeatable method for running the selection so that the answer survives contact with reality.

Anchor on requirements

The default selection process is run by the vendors, and most buyers never notice. You take the demos, you build a comparison grid from the capabilities you were shown, and by the third meeting your evaluation criteria are simply the union of what the vendors chose to present. The scoring feels objective. It is not. It is a beauty contest whose rules were written by the contestants, and the vendor with the most polished demo wins a game you did not know you were playing.

The correction is to write your requirements down before you see anything, and to separate two kinds. Table stakes are the requirements every serious candidate must meet: security controls, integration with your identity and systems of record, the core capability itself. They screen out; they do not differentiate, because everyone clears them. Differentiating criteria are the few dimensions on which your specific context makes one vendor genuinely better for you: latency in your workflow, quality on your domain's language, deployment in your regulatory geography, depth on the one integration that matters. Weight these before the demos, not after.

The point of anchoring first is not paperwork. It is to hold a fixed reference the vendor cannot move. When a demo dazzles on a capability that was not on your list, the anchored buyer asks a useful question: is this a criterion I under-weighted, or a feature I am being taught to want? Without the anchor, every impressive thing you are shown silently becomes a requirement, and the grid drifts toward whoever demos best.

Involve the people who own the receiving workflow and the data in setting the weights, not only the technologists who will integrate the tool. A requirement set written by the builders alone tends to over-index on the technically interesting and under-index on the operationally decisive, which is exactly the axis on which real programs succeed or fail.

Prove it on your data

Here is the step that decides the selection and the step that gets cut for time: run a proof of concept, or a head-to-head bake-off, on your data and your tasks. Not the vendor's curated corpus, not the sample the sales engineer has rehearsed against a hundred times, but a representative slice of your real inputs, with the messiness, the ambiguity, and the edge cases your production traffic actually contains. A vendor demo is a controlled experiment designed to produce one result. Your data is the only test whose outcome you have not already been told.

On demo theater: every vendor demo you will ever see has been tuned until it works. The corpus is clean, the questions are known, the failure cases are quietly out of frame. This is not deception so much as selection bias with a sales incentive, and it means a demo carries almost no information about how the product behaves on your inputs. The only demo worth trusting is the one running on data you supplied and questions you wrote.

Design the bake-off so the result means something. Give each finalist the same task set and the same data, define the scoring rubric in advance, and score with a method matched to the task rather than a vibe in the room, whether that is deterministic checks, semantic comparison against a reference, human rating with a clear protocol, or a calibrated LLM-as-judge (the trade-offs are laid out in the evaluation methods deep-dive). Cherry-picked wins prove nothing; sample size and variance matter here as much as anywhere.

The POC also surfaces what no slide will tell you: how the product fails, how legible those failures are, and how much engineering it takes to get from impressive to reliable on your domain. Two vendors can look identical in a demo and differ by an order of magnitude in the effort to reach production quality on your corpus. That gap is invisible until you have run your own data through both, which is precisely why the step you are tempted to skip is the one that pays for the whole evaluation.

The data terms

For a regulated buyer, the contract clauses governing data frequently matter more than any feature on the comparison grid. A feature gap you can engineer around or wait out; a data term that lets a vendor train on your inputs, or park them in the wrong jurisdiction, is a standing exposure that no amount of product quality offsets. Read these clauses before you fall in love with the demo, because they are far harder to renegotiate once you are committed.

Four terms carry most of the weight, and each deserves an explicit, written answer rather than a reassuring sentence in a sales call.

TermThe question to force onto paperWhy it bites later
Training rightsCan the vendor train or improve its models on your inputs and outputs, and is opt-out the default or an upsell?Your proprietary data can leak into a model other customers use; the default is often permissive.
Data residencyIn which jurisdictions is data processed and stored, and can you pin it to a required region?Cross-border processing can breach data-protection law regardless of product quality.
RetentionHow long is data kept, for what purpose, and can you require deletion on a defined clock?Indefinite retention widens the breach blast radius and complicates your own obligations.
No-train guaranteeIs the no-training commitment contractual and auditable, or a policy the vendor can revise?A policy is not a promise; only the contract survives a change of ownership or strategy.

These terms tie directly into the privacy and PII controls covered in the privacy deep-dive: residency, retention and training rights are the contractual surface of the same obligations your data-protection posture already owes. Treat the sub-processor list, the deletion mechanics, and the breach-notification window as part of this same review, not as legal boilerplate to be skimmed after the technical decision is made.

A test I apply: ask the vendor to put the no-train guarantee, the residency commitment, and the deletion clock into the contract, in writing, with an audit right. A vendor that will say it on a call but not sign it has told you which one it means. The gap between the marketing page and the redline is where the real terms live.

Lock-in and exit

Every vendor decision is also a decision about how expensive it will be to reverse. Lock-in is not a single thing; it accumulates in layers. There is the integration cost you pay to wire the product into your systems, the proprietary formats and APIs that make your work non-portable, the data that now lives in the vendor's shape, and the organizational muscle memory built around one product. None of these is visible in the pilot, and all of them compound into a switching cost that quietly transfers negotiating power to the incumbent at every renewal.

Price this at selection time, while you still have leverage, and negotiate the exit while the vendor still wants your signature. The clauses that matter are unglamorous: a data-portability right that lets you extract your data and any derived artifacts in a usable format, a defined exit and transition period, and clear terms on what happens to your data on termination. A vendor confident in its product rarely objects; resistance to exit terms is itself a signal worth reading.

Architecturally, the strongest defense against lock-in is to design a seam where the vendor plugs in rather than a weld. A model gateway (covered in the model gateways deep-dive) is the canonical example: route through an abstraction layer so that swapping a provider is a configuration change at one boundary rather than a rewrite scattered through your application.

      your application
             |
   [ gateway / abstraction seam ]   <- one swap point
       |         |         |
   vendor A   vendor B   vendor C    <- interchangeable
A gateway turns a vendor swap into a config change at a single seam.

The seam is not free, and it does not abstract away the deep integrations, the fine-tuned artifacts, or the workflow habits. But it converts the most volatile part of the stack, the model and its provider, into something you can re-shop as the market moves, which is exactly the freedom a young market makes valuable. This is the orchestrate posture from the build-versus-buy-versus-orchestrate framing applied as a hedge rather than a strategy.

Vendor viability

The AI vendor market is young, crowded, and churning. Startups pivot, get acquired, run out of runway, or are absorbed into a platform whose priorities are not yours. That volatility makes vendor viability a first-class evaluation criterion rather than diligence boilerplate: you are not just buying a product, you are betting the vendor will still be around, still investing, and still solvent long enough to be worth the integration you are about to pay for.

Four dimensions carry the viability judgment, and they deserve explicit weight in the scoring rather than a footnote after the feature grid. The weights below are illustrative; set your own by how much each risk would actually hurt in your context.

CriterionWhat to probeIllustrative weight
RoadmapDirection and credibility of the plan; does it move toward or away from your needs?High
Financial healthFunding stage, runway, revenue base, ownership; likelihood of surviving the next cycle.High
Security postureCertifications, sub-processor hygiene, breach history, independent audit.High
Support & SLAsResponse commitments, escalation path, versioning and deprecation policy.Medium

Financial health is the awkward one to ask about and the most consequential to skip. A brilliant product from a vendor with nine months of runway is a migration project you have not scheduled yet. Ask about funding, ownership and revenue base directly; a healthy vendor answers, and evasion is data. Security posture deserves evidence rather than assurances: certifications, a clean sub-processor list, and a willingness to be audited, all of which connect back to the data terms in the section above.

Weight viability against the reversibility you built in the previous section. If the abstraction seam makes a swap cheap, you can tolerate a riskier but stronger vendor; if switching is expensive, viability risk should dominate the score. The two sections are one judgment: how good the vendor is, discounted by how likely it is to survive and how hard it would be to leave.

The architect view

Strip the process to its load-bearing moves and vendor selection is not complicated, only easy to do badly under sales pressure. The failures are predictable: letting demos set the criteria, skipping the proof on real data, treating the contract as legal cleanup after the technical decision, and assuming the vendor you pick today is the vendor you will still want in three years. Each has the same antidote, which is to run the selection as procurement, on your terms, with the durable risks weighted ahead of the features.

  1. Anchor before the demos. Write your requirements and, crucially, your few differentiating criteria first, and weight them before any vendor frames your thinking.
  2. Prove it on your own data. Run a POC or bake-off on representative inputs with a scoring rubric fixed in advance. Never let the vendor's demo stand in for this.
  3. Win the data terms. Get training rights, residency, retention and a no-train guarantee into the contract, with an audit right, not just onto a call.
  4. Price and negotiate the exit. Cost the lock-in, secure data-portability and transition clauses, and keep a gateway as an abstraction seam so a swap stays a config change.
  5. Weight viability, then re-evaluate. Score roadmap, finances, security and support; then revisit the whole decision as a young market moves, because today's best vendor is a dated judgment by the next renewal.

The quiet opinion underneath all of it: in AI procurement the features are the most visible and least durable part of the decision, and the data terms, the exit path, and the vendor's survival are the least visible and most durable. A process that inverts its attention to match, spending its scrutiny where the risk actually lives, is the difference between a vendor choice you defend later and one you spend two years unwinding. Benchmark your own shortlist against this discipline with the RFP Vendor Comparison tool in the Lab.

← Value Realization: Proving AI ROI ALL OF STRATEGY AI Maturity Assessment: Knowing Where You Stand →