Atlas / GOVERN & LEAD / Strategy / Supplier Risk
DEEP-DIVE · STRATEGY

AI Supplier Risk: Designing for a Model You Might Lose

In 2026 a frontier model was switched off overnight by an export-control order, a lab paused training on safety grounds, and outage days multiplied. Intelligence is now a concentrated supply chain. How to tier workloads by dependency, what real portability costs, and the moves that keep operations running.

TL;DR
  • A model an operation depends on can be paused, repriced, withdrawn or silently changed on a timescale of weeks, and 2026 supplied examples of each: an export-control order that took two frontier models offline for every customer for two and a half weeks, a safety pause in frontier training, and a sharp rise in outage days across major assistants.
  • Not every workload needs portability. Tier them by criticality and substitutability, give the critical tier a tested secondary path, and accept the risk deliberately for the rest. Naming a backup provider is not a plan; moving a real workload is.
  • Portability is bought with portable assets, prompts, skills, tool definitions and above all evaluation suites, plus a gateway that can route between providers and contracts that cover deprecation, price changes, data export and legally mandated suspension.

Pause, reprice, withdraw

Enterprises spent 2024 and 2025 embedding language models into customer service, document processing, code review and financial analysis. By 2026 many of those workloads were load-bearing, and the question changed from which model is best to what happens when the model is not there. Four failure modes cover the ground:

Failure mode2026 exampleTypical noticeFirst defense
PauseOutages: one analysis counted 51 high-impact disruption days across major assistants in Q1 2026, against 6 a year earlierNoneFailover to a validated secondary
WithdrawA June 12 US export-control directive led a provider to disable two frontier models for all customers until June 30None to monthsPortable assets and a tested migration path
RepricePrice, rate-limit, terms and data-use changes; scheduled model retirementsDays to monthsRouting, cost ceilings, contractual price protection
DriftProvider-side updates that change accuracy, tone or safety behavior while your code is untouchedOften noneVersion pinning and continuous evaluation

A fifth pressure sits upstream of all four. In August 2026 OpenAI paused reinforcement-learning training of its newest models for two weeks and held its largest planned training run while it hardened safeguards. A training pause is not an outage, but it is a reminder that the arrival of the next capability tier, which many roadmaps quietly assume, is also a supplier decision. The working rule is blunt: if the model an operation runs on can disappear within weeks, the operation must be able to run on another one, or the business must have explicitly decided that it need not.

A concentrated supply chain

The structure of the market makes this a systemic issue rather than a vendor-management detail. A handful of labs supply the frontier models that most enterprise workloads depend on, and three hyperscale clouds supply most of the compute underneath them. Dependencies are also hidden: applications that look independent often share the same base model, the same cloud region or the same accelerator supply, so one event can fail many of them at once.

Executives know the exposure is real. In a mid-2026 IBM survey, 71 percent said switching their primary AI vendor would be difficult, and 81 percent said a seven-day outage would cause severe or critical disruption. Supervisors have noticed too. In September 2026 Moody's warned that lenders and insurers were consolidating their AI onto a handful of model builders and cloud platforms, that vendor updates can change a service's behavior while a bank's own code is untouched, and that naming a backup provider does not make a workload portable. The Financial Stability Board and the Bank of England have both identified supplier dependency as a channel through which AI could amplify systemic risk, and EU financial entities already manage ICT third-party concentration as a supervised risk under DORA.

Tier workloads by dependency

Portability is expensive, so spend it where it matters. Two questions place each workload: how critical is it, measured by how long the business can tolerate its absence, and how substitutable is the model underneath it.

                         SUBSTITUTABILITY
                    easy                  hard
               +---------------------+---------------------+
   CRITICAL    | TIER 1              | TIER 1+             |
   (hours)     | active secondary,   | invest to make it   |
               | tested failover     | easy, or keep a     |
               |                     | non-AI fallback     |
               +---------------------+---------------------+
   IMPORTANT   | TIER 2              | TIER 2              |
   (days)      | warm standby        | documented migration|
               |                     | plus annual drill   |
               +---------------------+---------------------+
   CONVENIENT  | TIER 3              | TIER 3              |
   (weeks)     | accept the risk     | accept, and watch   |
               +---------------------+---------------------+
Tier by tolerance for absence and by how hard the model is to replace. Only the top row justifies a live secondary; the bottom row should be an explicit, recorded acceptance of risk.

The uncomfortable quadrant is critical and hard to substitute: a customer-facing workflow tuned so closely to one model, or dependent on a capability only one model has, that nothing else passes its evaluations. The options there are to invest until substitution is possible, often by restructuring the task so a smaller or different model can handle most of it, or to keep a degraded non-AI path, such as rules or human handling, that the business can survive on for a while.

What portability really costs

Swapping models breaks more than an endpoint URL. Prompts and guides are tuned to one model's habits. Tool-calling and structured-output behavior differ in small ways that matter. Context limits, refusal behavior and latency profiles change. Fine-tuned adapters do not transfer, as the fine-tuning deep-dive explains. And embeddings are the trap most teams discover last: vectors from one embedding model are meaningless to another, so switching means re-embedding the whole corpus.

The instrument that makes switching possible is the evaluation suite. You can move to a substitute only as fast as you can prove it works on your tasks, which makes the golden sets and judges in the evals deep-dive the core portability asset, not a quality nicety. Around them sit the other portable assets: prompts, skills and guides stored as files; tool access through MCP rather than vendor-specific plugins; source text and a working re-embedding pipeline kept alongside every vector index; and the training data behind any fine-tune.

Naming a backup is not a plan. A secondary provider that has never carried production traffic is a hypothesis. The test, as Moody's put it to banks, is to actually move a critical workload across and see how the substitute behaves in production. Do it before you need to.

There is also a portability tax. Abstraction layers that hide every provider difference drift toward the lowest common denominator and forfeit the features that made a model worth choosing. Abstract the interface, not the capability: route through a common gateway, keep provider-specific features behind small adapters, and accept that the secondary path may be somewhat worse, as long as it is good enough for the tier.

Resilience patterns

The patterns are familiar from cloud resilience, adapted for models:

Contracts and governance

Architecture handles the outage; contracts handle the rest. The clauses worth negotiating, in roughly descending order of value: advance notice periods for model deprecation and a commitment to keep pinned versions available for a defined time; notification of material model updates; price protection or caps for the contract term; export rights for logs, fine-tuning data and anything else you would need to move; service levels with meaningful remedies, while remembering that credits do not run operations; exit assistance; and explicit treatment of legally mandated suspension, such as export controls or sanctions, including what the provider will do and how quickly it will tell you. The vendor selection deep-dive places these alongside the rest of the procurement checklist.

On the governance side, AI supplier risk belongs on the enterprise risk register, owned and reported like cloud concentration risk. Tier 1 dependencies deserve board-level visibility. And the work should connect to existing third-party risk management and to model risk management, because a supplier's model change is, in effect, a change to your model, a point the model risk deep-dive develops.

Credits do not run operations. A service-level agreement that pays you back a fraction of a monthly fee for a week-long outage is an accounting mechanism, not a resilience plan. Negotiate for notice, versions and exit rights, which change what you can do, rather than for credits, which only change what you are paid.

The architect view

Intelligence has become a critical supply, and it should be managed the way enterprises learned to manage cloud concentration a decade ago: deliberately, by tier, with tested exits for what matters and explicit acceptance for what does not. The year 2026 removed any doubt that the risk is real; what remains is the discipline to price it.

Four commitments make that concrete. Build the dependency inventory and tier every AI workload by criticality and substitutability. Keep prompts, skills, tool definitions and evaluation suites as portable assets, with evaluations treated as the switching instrument. Give Tier 1 workloads a live, scored secondary path, and an open-weight or non-AI fallback where the business requires one. And write contracts that cover pause, reprice and withdrawal, not just availability.

The model you depend on most is the one you should be most prepared to lose. Preparing costs a little every quarter; not preparing costs everything in the one week it matters.

← Driving AI Adoption: The Change Management Problem ALL OF STRATEGY