- Two organizations with the same models and the same cloud will get very different results. The operating model, which cannot be bought, is the actual differentiating asset.
- A Center of Excellence is a lifecycle stage, not a destination: incubate scarce skills, enable the business with paved roads, then dissolve delivery into the line.
- Fund AI capabilities as products with persistent teams and run-state budgets from day one; project funding that expires at the demo is how pilots die.
Why operating models decide AI outcomes
Most enterprise AI strategies I am asked to review are technology documents. They name models, platforms, vendors, and reference architectures, and they are often competent. Then the pilots run, the demos land well, and a year later the portfolio review shows fifteen proofs of concept and two systems in production, neither owned by anyone in particular. The technology did not fail. The organization around it was never designed.
Pilots die in a predictable gap between IT and the business, and three failure modes account for most of the losses. First, unclear ownership: the data science team built it, a business unit sponsored it, and nobody signed up to run it, so the prompts rot, the data drifts, and the first incident has no owner. Second, project funding that expires: the budget was scoped to deliver a demo, and it did, at exactly the moment the system needed run-state money for monitoring, evals, retraining, and support. Third, no defined path from demo to run-state: security review, data agreements, on-call rotations, and integration work were deferred to "phase two," and phase two had no sponsor.
None of these are technical problems, and no amount of model quality fixes them. They are operating model problems: decisions about who does the work, who pays for it, and who decides. My working observation, after enough of these reviews, is that two enterprises with near-identical technology stacks routinely end up with outcomes that differ by an order of magnitude, and the difference traces back to these decisions. That is why this article sits on the Strategy shelf rather than in Platforms.
The three archetypes
Three archetypes cover most of the real estate. Centralized: one AI team owns the platform, the use-case portfolio, and delivery. Federated: each business unit builds its own AI capability, with light central coordination at most. Hub-and-spoke: a central hub owns the platform, standards, and the scarcest expertise, while spokes embedded in the business own use cases and delivery.
| Archetype | Strengths | Weaknesses | Fits best when |
|---|---|---|---|
| Centralized | Consistency; concentrates scarce talent; one governance surface; standards set fast | Becomes a bottleneck at scale; distant from domain context; business units disengage and wait in line | Early maturity, scarce skills, high regulatory stakes |
| Federated | Speed and domain intimacy; genuine business ownership of outcomes | Duplicated platforms and vendor contracts; inconsistent risk posture; quality varies wildly by unit | High AI literacy across units and a strong shared platform already exist |
| Hub-and-spoke | Shared platform economics with domain-led delivery; talent has a home and a field position | Interface friction between hub and spokes; the hub can drift into gatekeeping if roles are not explicit | Mid-to-late maturity, multiple units with real demand |
The honest reading of that table is that every archetype fails in the direction of its strength. Centralized models buy consistency and pay for it in throughput. Federated models buy speed and pay for it in duplication and uneven risk. Hub-and-spoke looks like the answer on paper, and often is, but it demands the most deliberate role design, because an underspecified hub degenerates into either a helpdesk or a tollbooth.
The debates I sit through treat this as an ideological choice, as if one archetype were simply correct. It is not. The right answer is a function of maturity: how scarce your skills are, how literate your business units are, and how much platform you have already built. An honest maturity assessment answers the archetype question faster than any org-design workshop.
The CoE as a lifecycle, not a destination
The Center of Excellence question generates more heat than any other in this space, and most of the heat comes from treating the CoE as a permanent org unit that is either good or bad. The more useful frame is temporal: a CoE is a lifecycle stage, and the design question is not whether to have one but how it should change shape over time.
INCUBATE ENABLE DISSOLVE +--------------+ +--------------+ +--------------+ | central team | | central team | | line teams | | delivers the | ---> | paves roads, | ---> | own AI as | | first wins | | trains, sets | | normal work | | itself | | standards | | | +--------------+ +--------------+ +--------------+ scarce skills delivery shifts center retains concentrated to the business platform + risk
Incubate: when skills are genuinely scarce, concentrate them. The central team delivers the first production wins itself, end to end, because at this stage the fastest way to learn what works in your enterprise is to do the work. Enable: the team's job shifts from delivering use cases to making delivery easy for others: paved-road templates, an evaluated model catalog, training curricula, reusable governance artifacts. Dissolve into the line: AI stops being a special activity done by special people and becomes how underwriting, service, and supply chain do their jobs. What remains central is the platform and the risk function, not delivery.
The anti-pattern is the CoE that never advances past stage one: a permanent gatekeeper through which every use case, model choice, and prompt change must pass. It starts as quality control and ends as a queue. Business units learn to route around it with shadow AI, which is the worst of both worlds: no consistency and no visibility. A CoE that measures its success by how much it does, rather than by how little the business needs it, has inverted its own purpose.
Fund products, not projects
Funding model is the operating model decision that most reliably predicts whether an AI system survives its first year, and it is usually made by default rather than by design. The default is project funding: a fixed budget, a delivery date, a steering committee, and a handover. That framing is lethal to AI systems, because an AI capability is never finished. Models get upgraded, prompts and evals need continuous tuning, data drifts, usage reveals new failure modes. A project ends; the system's needs do not.
The handover is where the damage concentrates. The delivery team that holds all the context disbands, the receiving team inherited neither the eval suite nor the design rationale, and the run budget was an afterthought line item. Within two quarters the system is frozen: nobody dares change it, so it quietly degrades until someone proposes a new project to replace it. I have watched this cycle consume entire AI budgets while producing nothing durable.
The alternative is product funding: a persistent, cross-functional team that owns an AI capability over its life, measured on outcome metrics (cases resolved, cycle time, loss rates) rather than delivery milestones. The same logic applies one level down: the shared platform should be run as a product with internal customers, adoption targets, and a roadmap, not as a cost center that fields tickets. This position is argued more than it is proven, so I will frame it as advisor judgment, but it is judgment I hold with confidence: every durable AI capability I have seen up close was product-funded, and most of the abandoned ones were projects.
Talent topology
Operating models are ultimately arrangements of people, and AI work needs a triad of roles that rarely exist in one person. AI architects make the system-level calls: model selection, retrieval design, integration patterns, the build-versus-buy line. AI engineers ship and run the systems: pipelines, evals, monitoring, the unglamorous bulk of the work. Domain translators sit in the business and convert process knowledge into use cases, requirements, and honest acceptance criteria. Enterprises consistently over-hire the first two and under-invest in the third, then wonder why technically sound systems solve the wrong problems.
Where these people sit should follow scarcity. Early on, architects and senior engineers are rare enough that spreading them across business units wastes them; concentrate them centrally, where they compound each other. Translators are the opposite: they only work embedded in the domain, so grow them in the line from day one, deliberately, through rotations into the central team and structured training rather than a one-off literacy webinar. Federated literacy is built, not announced.
The build-hire-partner question deserves more honesty than it usually gets:
- Build (reskill your own people) is slow but durable, and it is the only source of translators, because domain knowledge cannot be hired in.
- Hire works for engineers and architects, but the market is competitive, and senior AI talent evaluates your platform and autonomy the way you evaluate their resume.
- Partner buys speed and pattern knowledge from firms that have seen many deployments, at the price of context walking out the door at contract end. Structure every partnership with explicit skill-transfer obligations, or you are renting capability, not acquiring it.
A defensible default: partner to incubate, hire to run, build to scale.
Decision rights and governance interfaces
When AI initiatives stall, the proximate cause is usually an unmade decision that nobody was empowered to make. Who approves a new use case? Who can choose or swap a model? At what spend threshold does procurement engage? Who signs off that a system touching customers is safe to ship? In most enterprises these questions have no standing answer, so each one convenes a meeting, and the meetings compound into the real cost of governance.
The fix is not a bigger committee. It is treating governance as an interface: a published contract that tells a delivery team what they can do freely, what needs review, and how long review takes. The critical design move is splitting the routine from the novel. A use case built on the paved road (approved models, standard data classes, established patterns) should clear governance in days through a checklist, not a board. Intake review, with real scrutiny from risk, legal, and security, should be reserved for what is genuinely novel: new data categories, new model providers, customer-facing autonomy, regulated decisions. Route everything through the same gate and you get a queue; the queue creates shadow AI; shadow AI creates the incidents the gate existed to prevent.
A lightweight RACI is enough if it is published and enforced:
- Use-case approval: business unit owner decides, platform team and risk are consulted, portfolio council is informed.
- Model and vendor choices: platform team decides within an approved catalog; additions to the catalog go through architecture and security review.
- Spend: product teams operate freely under a threshold; above it, standard investment governance applies.
- Risk acceptance: the risk function decides for high-stakes classes, on the classification logic covered in Governance.
One page, four rows, publicly known. That beats any charter deck.
An evolution path
Everything above compresses into one principle: the operating model should be staged to maturity, and it should be expected to change. The design failure is not picking the wrong archetype; it is picking any archetype and declaring it permanent.
- Centralize while scarce. Concentrate talent, deliver the first wins centrally, stand up the platform, and keep governance simple because volume is low. Fund the central team as a product from the start, even now.
- Federate as literacy grows. Shift the central team toward enablement, push delivery into business units behind paved roads, grow translators in the line, and split governance into intake versus fast-path.
- Re-centralize only platforms and risk. Delivery lives in the line for good; the durable central functions are the platform product, the model catalog, the eval and observability stack, and risk oversight. Nothing else needs to be central anymore.
Timelines vary too much to be worth stating; I have seen the first transition take a year and I have seen it take four. The signals matter more than the calendar. It is time to shift from stage one to stage two when the central backlog is growing faster than throughput, when business units start hiring their own AI people regardless of policy, and when the central team's calendar fills with support requests rather than delivery. It is time for stage three when paved-road use cases outnumber novel ones, when spokes ship without the hub noticing, and when governance findings cluster in the long tail rather than the mainstream.
Two failure signals warrant a step backward: rising incident rates in federated teams, or platform fragmentation as units quietly rebuild shared components. Neither is a defeat; recentralizing a specific capability for a period is the model working as designed. Benchmark where you stand with the maturity assessment in the Lab, and treat the answer as this year's, not the final one. The operating model is not a decision you make once. It is a capability you run.