What is an AI Center of Excellence?
A central team that holds the shared parts of AI work so every business unit is not solving the same problems separately. That covers the platform, the standards, the reusable components and the evaluation harness, and the interesting question is how much authority it holds over the rest.
The term is used loosely enough that two organisations can both have one and mean entirely different things. What separates them is not the label but the answer to a single question: does the CoE build the AI, or does it make it possible for others to build it? Almost every disagreement about a CoE traces back to that.
A note on scope, since two questions hide behind the same phrase. This page covers what a CoE is, which operating model fits, and how to measure one. Designing, staffing and standing one up for your organisation is a services question, covered at establish AI centers of excellence.
What are the three operating models?
Centralised, federated and hybrid, and they differ on where the work happens rather than on where the org chart puts people. Most organisations start centralised because that is the only option when the skill is scarce, and most end up hybrid whether or not anybody decided to.
The distinction that matters is delivery, not reporting lines.
| Model | Who builds | Where it fails |
|---|---|---|
| Centralised | The CoE builds for everyone | A queue forms, and the CoE never learns any domain deeply enough |
| Federated | Business units build, the CoE supplies the platform and standards | Requires capability in the business units that frequently is not there yet |
| Hybrid | The CoE builds the first case in a domain, then hands over | The handover never happens unless somebody is measured on it |
The third row is where most organisations land, and the failure noted against it is the one to watch. Handover is the step that gets postponed indefinitely, because the CoE is busy and the receiving team is not ready, and both of those things are true every quarter.
How do you decide which AI model fits in for a requirement?
Ask which resource you are short of, the ability to build, or the knowledge of what to build. Engineering skill moves between teams easily, so a central group makes sense when that is your real constraint. But judgment, like understanding how underwriting works, does not move that way. No central group will ever build up enough of that kind of deep, specific knowledge on its own.
The CoE should own everything identical across use cases and nothing specific to one.
This is the decision rule, and it is more useful than any maturity model because it is answerable this afternoon.
- Scarce skill, abundant context. Few people can build, and the requirements are clear once explained. Centralise. The CoE can learn enough of each domain to be useful.
- Abundant skill, scarce context. Plenty of engineers, and what counts as a correct outcome lives in the heads of people who have done the job for fifteen years. Federate, because a central team will never know enough and will produce things that pass review and fail in practice.
- Both scarce. Hybrid, deliberately. The CoE builds the first case in a domain alongside the people who own it, and the point of that first case is the handover rather than the delivery.
- Both abundant. You may not need a CoE at all, only shared infrastructure and a standard. This is rarer than teams think and worth testing before funding a team.
The second row is the one organisations get wrong most expensively. A central team can be given the domain knowledge or it can be given the domain, and only the second actually works. Requirements documents are how context is lost rather than how it is transferred, which is why the evaluation criteria discussed at agent evaluation are so hard for a central team to write alone.
What should a CoE own, and never own?
It should own everything identical across use cases and nothing specific to one. The moment a CoE owns a use case it has become a delivery team with a waiting list, and the shared assets that justified its existence stop being maintained.
The split is cleaner than most operating model discussions suggest.
- Owns the platform and the control plane. Registry, identity scoping, audit, cost attribution, promotion gates. Identical for every team and expensive for each to build, which is the argument at build vs buy.
- Owns the standards and the patterns. How an agent is registered, what must be traced, which approval gates apply. Enforced through the platform rather than through documents, per agent governance.
- Never owns the use case. The business unit does, because it carries the consequence when the agent gets something wrong.
- Never owns what counts as correct. That is domain knowledge, and a CoE writing evaluation criteria for a process it does not run is guessing politely.
Why do CoEs become bottlenecks?
Because success generates demand faster than a central team can absorb it. A CoE that delivers well becomes the place every good idea is sent, and the queue in front of it grows at the rate the organisation gets enthusiastic rather than at the rate the team can hire.
Four things happen next, and the last one is the expensive one.
- 01
The queue becomes the constraintDelivery capacity stops being the limit and prioritisation does. What follows: the CoE spends more time deciding what not to do than building, which is a governance role nobody designed it for.
- 02
Standards get skipped to clear the queuePressure to deliver erodes exactly the practices the CoE exists to hold. What follows: the central team becomes the largest source of exceptions to its own standard.
- 03
The shared assets stop being maintainedPlatform work is never urgent against a named delivery deadline. What follows: the reusable components decay, which removes the reason the CoE existed.
- 04
Teams route around itBusiness units start building without telling anybody, because the approved route is slower than the unapproved one. What follows: the estate the CoE was created to govern is now partly invisible to it. See shadow AI.
The pattern is the same one that breaks approval boards, arriving through a different door. A central function cannot scale linearly against demand that compounds, so the choice is between deliberately federating and being federated around. Only one of those leaves you with visibility.
How do you measure the success of an AI CoE?
Count the work it did not touch. Any team scored on use cases delivered is effectively being funded to remain in the critical path, and it will remain there, because that is what the scoring rewards. The figure worth reporting is how much AI reached production across the business with the central team nowhere near it.
Five measures, and the first one changes behaviour more than the other four together.
-
Work delivered by others on your platform
Agents built by business units, using the CoE's platform and standards, without the CoE building them. This is the enabling metric, and it rises only when the CoE invests in making itself unnecessary.
-
Time from idea to first working version
Measured for the requesting team, including the wait. A CoE that builds quickly but starts in eleven weeks is slower than it looks from inside.
-
Reuse rate of shared components
How often a new agent uses existing patterns rather than starting fresh. Falling reuse is the early signal that the shared assets have decayed.
-
Share of the estate that is registered
Agents in the registry against agents discovered. A widening gap means teams are routing around the approved path, which is a measurement of the CoE's usefulness rather than of anybody's compliance.
-
Agents retired
Almost never tracked, and the only metric that acknowledges an estate has a lifecycle. An organisation that has retired nothing has not yet had the ownership conversation.
Settle one structural question early, because it gets awkward if you leave it for later. When you create the AI Center of Excellence, write down what would have to be true for it to eventually be disbanded.
This is not a threat. It is a design choice. Here is why it matters. A team that knows what ‘finished’ looks like naturally advances toward handing things off to others. But a team with no such definition of ‘done’ will start optimizing to keep itself around. That is not a criticism of that team. It is just what any of us would do in the same position.