What is an AI platform?
AI platform usually means the shared layer your AI runs on. That definition is broad enough to describe four unrelated products. Vendors did not plan the confusion. The category formed quickly and the language never settled. Resolving this definition is always the first step in any serious evaluation.
The practical consequence appears early and is usually misdiagnosed. A shortlist is assembled, the demos look incomparable, and the team concludes the market is immature. More often the shortlist contains products from two or three different categories, each answering a question the others do not address at all.
A note on scope, since two different questions hide behind the same phrase. This page covers what an AI platform is and how to evaluate one. Who designs, builds and runs it for you is a services question, covered at AI platform engineering services.
Which are the four things referred as one under AI Platform?
Infrastructure, model serving, model lifecycle and agent operations. They sit at different heights in the stack, solve different problems and are frequently bought from different vendors, so a comparison table with all four in it will compare things that do not overlap.
Sorting your shortlist into these four is usually a ten-minute exercise that changes the conversation.
| Category | The question it answers |
|---|---|
| Infrastructure | Where does this run, and on what hardware. Largely a cloud decision you have already taken |
| Model serving | How do applications reach models, with what keys and what failover. See model routing |
| Model lifecycle | How does a model get trained, deployed and monitored. The MLOps inheritance, and the category most often mistaken for the next one |
| Agent operations | What agents exist, what may each do, what did each do. The layer covered at agent governance |
The two marked rows are where evaluations go wrong, because their vocabularies are nearly identical and their abstractions are not. Both talk about versions, deployment, monitoring and governance, and mean different objects by every one of those words.
Why do MLOps-era platforms fit agents badly?
Because they were designed around a model you train and own. That assumption shaped every abstraction in them, and an agent calling somebody else's model through an API breaks it, so the concepts do not translate even when the words survive.
Four inherited assumptions, each of which stops holding.
- The unit is a model version. For an agent the unit is a configuration: a prompt, a tool set, a permission scope. Those change weekly and independently of any model.
- Monitoring means drift. Statistical drift in inputs is a real signal for a trained model. An agent fails by taking a reasonable-looking wrong action, which no drift metric detects. See AI agent observability.
- Testing means a held-out set. There is no labelled test set for whether an agent handled a refund request well, which is the problem described at agent evaluation.
- Governance means model documentation. Model cards and lineage are genuine artefacts about something you trained. They say nothing about which systems an agent may reach or on whose authority.
Which layers are commodity, and which are not?
The lower you go, the more commoditised it is, and the buying decision should follow that gradient rather than treating all four layers as equally consequential. Most organisations have already settled the bottom two without calling it a platform decision.
Where effort is worth spending, from the bottom up.
- 01
Infrastructure: settledDecided by your cloud relationship, usually years ago. What to do: confirm it meets any residency constraint and move on. Re-opening it to buy an AI platform is the tail wagging the dog.
- 02
Model serving: convergingGateways and multi-provider access have become close to interchangeable. What to do: treat as a commodity and prioritise the ability to change provider, since that is the capability you will actually use.
- 03
Model lifecycle: only if you trainGenuinely valuable if you build your own models, and largely irrelevant if you consume them. What to do: be honest about which you do, because this is where budget goes to serve a workload some organisations do not have.
- 04
Agent operations: the open decisionThe least commoditised and the one that scales with your agent population. What to do: spend the evaluation effort here, since this is where the difference between products is real.
Whether you buy this layer at all is a separate argument with a genuine case on both sides, set out at build vs buy for agent platforms. What matters here is narrower: the platform question is really a question about the fourth layer, and the first three tend to consume the meeting.
What does an agentic workload need?
You need less than most vendor pitches suggest. One requirement rarely appears in those pitches. Your needs fall into two groups. The first group covers what you can add later using tools you already own. The second group is small. Those items must exist on day one, because you cannot add them afterward.
Three things belong in the second group, and the third is the one to test hardest.
- Discovery of agents you did not create through it. A platform that only knows about agents built inside it will miss whatever appears in the tools your business already licenses, which is the mechanism described at shadow AI.
- Identity that reaches the model call. If the requesting person is lost at the first service boundary, neither the audit question nor the cost question is answerable afterwards. See LLM cost attribution.
- An exit. Whether your agents, their configuration and their history can leave. This is rarely on a requirements list and is the one that determines how much the other decisions cost you later.
Everything else on a typical requirements document is either assemblable or already present somewhere in your estate. The full control plane inventory, with what each component takes to build, sits at build vs buy rather than being repeated here.
How do you evaluate an AI Platform?
By testing operation rather than creation. Every platform in this category demos well, because the demo shows an agent being built in minutes and building an agent was never the difficult part. What separates products is what happens on day ninety.
Five tests, none of which appear in a standard demo script.
-
Sort the shortlist into the four categories first
Before any demo. If two vendors turn out to be in different categories, you are not choosing between them and the evaluation should be split.
-
Ask what the primary object is
Model or agent. The answer tells you which heritage the product carries and therefore which concepts will be second-class, regardless of what the feature list says.
-
Bring an agent you built elsewhere
Rather than building one in the demo. How much has to change to run it there is the portability answer, and it is far more informative than a greenfield build.
-
Ask it to find something it did not create
An agent running in a platform you already license. Discovery across tools the vendor does not own is the capability most often assumed and least often present.
-
Ask for the export
Agents, configuration, traces and cost history, in a usable format. A vendor with a good answer will say so immediately, and the hesitation is itself the finding.
One framing to carry into the room. You are buying an operating model, not a set of features, and the features that get demonstrated are drawn from the half of the work that was never hard. The useful demo is a boring one: show me a hundred agents, tell me who owns the twelfth, and stop it while I watch.