All glossary terms
A Architecture Procurement

AI platform

Four product categories are sold under the name AI platform. Most shortlists mix at least two, such as model and agent builders. Those vendors solve different problems and never compete. The evaluation stalls and nobody notices.

Definition

An AI platform is the shared layer an organisation runs its AI on, covering some combination of compute, model access, lifecycle tooling and governance. The term spans four distinct product categories, which is why two vendors can both accurately call themselves an AI platform and have almost nothing in common.

What is an AI platform?

AI platform usually means the shared layer your AI runs on. That definition is broad enough to describe four unrelated products. Vendors did not plan the confusion. The category formed quickly and the language never settled. Resolving this definition is always the first step in any serious evaluation.

The practical consequence appears early and is usually misdiagnosed. A shortlist is assembled, the demos look incomparable, and the team concludes the market is immature. More often the shortlist contains products from two or three different categories, each answering a question the others do not address at all.

A note on scope, since two different questions hide behind the same phrase. This page covers what an AI platform is and how to evaluate one. Who designs, builds and runs it for you is a services question, covered at AI platform engineering services.

Which are the four things referred as one under AI Platform?

Infrastructure, model serving, model lifecycle and agent operations. They sit at different heights in the stack, solve different problems and are frequently bought from different vendors, so a comparison table with all four in it will compare things that do not overlap.

Sorting your shortlist into these four is usually a ten-minute exercise that changes the conversation.

Four product categories stacked, all commonly called an AI platform. At the base, infrastructure platforms providing compute, GPUs and orchestration. Above that, model serving platforms providing hosted model access, gateways and routing. Above that, model lifecycle platforms carrying the MLOps inheritance, covering training, deployment and drift monitoring. At the top, agent operations platforms covering registry, permissions, audit and cost attribution. A note states that most shortlists mix two or more of these and that each answers a question the others do not.
The four categories sold as AI platforms and what each actually answers
Category The question it answers
Infrastructure Where does this run, and on what hardware. Largely a cloud decision you have already taken
Model serving How do applications reach models, with what keys and what failover. See model routing
Model lifecycle How does a model get trained, deployed and monitored. The MLOps inheritance, and the category most often mistaken for the next one
Agent operations What agents exist, what may each do, what did each do. The layer covered at agent governance

The two marked rows are where evaluations go wrong, because their vocabularies are nearly identical and their abstractions are not. Both talk about versions, deployment, monitoring and governance, and mean different objects by every one of those words.

Why do MLOps-era platforms fit agents badly?

Because they were designed around a model you train and own. That assumption shaped every abstraction in them, and an agent calling somebody else's model through an API breaks it, so the concepts do not translate even when the words survive.

Four inherited assumptions, each of which stops holding.

  • The unit is a model version. For an agent the unit is a configuration: a prompt, a tool set, a permission scope. Those change weekly and independently of any model.
  • Monitoring means drift. Statistical drift in inputs is a real signal for a trained model. An agent fails by taking a reasonable-looking wrong action, which no drift metric detects. See AI agent observability.
  • Testing means a held-out set. There is no labelled test set for whether an agent handled a refund request well, which is the problem described at agent evaluation.
  • Governance means model documentation. Model cards and lineage are genuine artefacts about something you trained. They say nothing about which systems an agent may reach or on whose authority.

Which layers are commodity, and which are not?

The lower you go, the more commoditised it is, and the buying decision should follow that gradient rather than treating all four layers as equally consequential. Most organisations have already settled the bottom two without calling it a platform decision.

Where effort is worth spending, from the bottom up.

  1. 01
    Infrastructure: settledDecided by your cloud relationship, usually years ago. What to do: confirm it meets any residency constraint and move on. Re-opening it to buy an AI platform is the tail wagging the dog.
  2. 02
    Model serving: convergingGateways and multi-provider access have become close to interchangeable. What to do: treat as a commodity and prioritise the ability to change provider, since that is the capability you will actually use.
  3. 03
    Model lifecycle: only if you trainGenuinely valuable if you build your own models, and largely irrelevant if you consume them. What to do: be honest about which you do, because this is where budget goes to serve a workload some organisations do not have.
  4. 04
    Agent operations: the open decisionThe least commoditised and the one that scales with your agent population. What to do: spend the evaluation effort here, since this is where the difference between products is real.

Whether you buy this layer at all is a separate argument with a genuine case on both sides, set out at build vs buy for agent platforms. What matters here is narrower: the platform question is really a question about the fourth layer, and the first three tend to consume the meeting.

What does an agentic workload need?

You need less than most vendor pitches suggest. One requirement rarely appears in those pitches. Your needs fall into two groups. The first group covers what you can add later using tools you already own. The second group is small. Those items must exist on day one, because you cannot add them afterward.

Three things belong in the second group, and the third is the one to test hardest.

  • Discovery of agents you did not create through it. A platform that only knows about agents built inside it will miss whatever appears in the tools your business already licenses, which is the mechanism described at shadow AI.
  • Identity that reaches the model call. If the requesting person is lost at the first service boundary, neither the audit question nor the cost question is answerable afterwards. See LLM cost attribution.
  • An exit. Whether your agents, their configuration and their history can leave. This is rarely on a requirements list and is the one that determines how much the other decisions cost you later.

Everything else on a typical requirements document is either assemblable or already present somewhere in your estate. The full control plane inventory, with what each component takes to build, sits at build vs buy rather than being repeated here.

How do you evaluate an AI Platform?

By testing operation rather than creation. Every platform in this category demos well, because the demo shows an agent being built in minutes and building an agent was never the difficult part. What separates products is what happens on day ninety.

Five tests, none of which appear in a standard demo script.

  1. Sort the shortlist into the four categories first

    Before any demo. If two vendors turn out to be in different categories, you are not choosing between them and the evaluation should be split.

  2. Ask what the primary object is

    Model or agent. The answer tells you which heritage the product carries and therefore which concepts will be second-class, regardless of what the feature list says.

  3. Bring an agent you built elsewhere

    Rather than building one in the demo. How much has to change to run it there is the portability answer, and it is far more informative than a greenfield build.

  4. Ask it to find something it did not create

    An agent running in a platform you already license. Discovery across tools the vendor does not own is the capability most often assumed and least often present.

  5. Ask for the export

    Agents, configuration, traces and cost history, in a usable format. A vendor with a good answer will say so immediately, and the hesitation is itself the finding.

One framing to carry into the room. You are buying an operating model, not a set of features, and the features that get demonstrated are drawn from the half of the work that was never hard. The useful demo is a boring one: show me a hundred agents, tell me who owns the twelfth, and stop it while I watch.

Frequently asked questions about AI platforms

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is an AI platform?

The shared layer an organisation runs its AI on, covering some mix of compute, model access, lifecycle tooling and governance. That definition is broad enough to be true of four unrelated products, which is an artefact of how fast the category formed. Resolving which category each vendor belongs to is the first thing any evaluation should do.

What are the four types of AI platform?

Infrastructure, model serving, model lifecycle and agent operations. Infrastructure answers where things run. Model serving answers how applications reach models. Model lifecycle answers how a model gets trained, deployed and monitored. Agent operations answers what agents exist, what each may do and what each did. They sit at different heights and are frequently bought from different vendors.

Why do AI platform evaluations stall?

Usually because the shortlist contains products from two or three different categories, each answering a question the others do not address. The demos look incomparable and the team concludes the market is immature. More often the market is fine and the shortlist is mixed, which a ten-minute sorting exercise resolves before anybody sits through a demo.

What is the difference between an MLOps platform and an agent platform?

What they assume you own. MLOps platforms were designed around a model you train, so their unit is a model version, monitoring means statistical drift, testing means a held-out set and governance means model documentation. An agent calling somebody else's model through an API breaks all four assumptions, even though both product categories use identical vocabulary for entirely different objects.

How can you tell which heritage a platform has?

Ask what its primary object is. If everything in the interface hangs off a model, it is a lifecycle platform with agent features attached, and the agent concepts will be second-class however convincing the demo. That is not a reason to reject it; it is a reason to know what you are buying, because your own team will fill the gap.

Which AI platform layers are commoditised?

The lower ones. Infrastructure was settled by your cloud relationship years ago, and re-opening it to buy an AI platform is the tail wagging the dog. Model serving has converged to the point where gateways are close to interchangeable, so the capability worth prioritising is the ability to change provider. Agent operations is the layer where product differences are still real.

Do you need a model lifecycle platform?

Only if you train your own models. It is genuinely valuable if you build them and largely irrelevant if you consume them through an API, and being honest about which you do is worth doing early. This is a common place for budget to go towards serving a workload the organisation does not actually have.

What must an AI platform provide for agents?

Three things that cannot be added later. Discovery of agents it did not create, since a platform that only knows its own will miss whatever appears in tools you already license. Identity that survives all the way to the model call, or neither the audit nor the cost question is answerable. And an exit, meaning your agents, configuration and history can leave.

Why is an AI platform demo misleading?

Because it shows an agent being built in minutes, and building an agent was never the difficult part. Every product in the category demos well on creation. What separates them is what happens on day ninety, when there are a hundred agents, several people have left, and somebody asks who owns the twelfth one and whether it can be stopped.

What should you actually test in an evaluation?

Operation rather than creation. Bring an agent you built elsewhere and see how much must change to run it there. Ask the platform to find something it did not create, in a tool you already license. Ask for an export of agents, configuration, traces and cost history. A vendor with a good answer says so immediately, and hesitation is itself the finding.

Is an AI platform the same as an agent platform?

No, though the terms are used interchangeably. An agent platform is one of the four things called an AI platform, sitting at the top of the stack and concerned with what agents exist, what each may do and what each did. The other three concern compute, model access and the training lifecycle, and an organisation can need any combination.

Who builds an AI platform for you?

That is a services question rather than a product one, and worth separating from this page. Defining what an AI platform is and evaluating which to buy is one exercise; designing, integrating and running one against your existing estate is another, usually involving different people and a different budget. Whether to build the agent operations layer at all is a third question again.

Ask for the boring demo
Show me a hundred agents and tell me who owns the twelfth

SERAA Cortex operates at the fourth layer: auto-discovery of agents across Gemini Enterprise, Azure AI Foundry, AWS Bedrock and Databricks AgentBricks into one registry with version lineage, prompts, keys and rules scoped per tenant, promotion through Dev, QA and Production, and an instant kill switch for any agent on any cloud.