Table of Contents
Pilot, Platform Team, or CoE: the Three AI Agent Org Designs That Work
Every company that runs AI agents at scale has settled into one of three org designs, or stalled because it tried to skip one. Each fits a headcount band and a governance reality, and each turns into a liability the moment you hold it past that band. Match the design to your engineer count and your compliance posture and the structure funds itself. Get the timing wrong and you get one of two failure modes: a billing shock nobody can explain, or two people writing policy docs while nobody ships.
Here are the three, each with its headcount band and mandate.
Embedded champions. One or two AI-fluent engineers spread across product teams, no central infrastructure. The right design for a pilot under roughly 15 engineers. Mandate: prove the technology works in this org.
Central platform team. A dedicated team of 2 to 8 that owns the agent infrastructure, the tooling, and the internal APIs. Right from about 15 to 150 engineers actively using AI tools. Mandate: make the technology safe and cheap for everyone else to use. In the Team Topologies sense, a platform team is "a grouping of other team types that provide a compelling internal product to accelerate delivery by Stream-aligned teams" per Team Topologies.
Center of Excellence. A named org with budget, headcount, a governance charter, and an advisory role to the business units. Right above 150 engineers, or wherever regulatory posture forces formal sign-off. Mandate: own the policy, not only the plumbing.
The decision table
Paste this into your next leadership doc. Read down the column that matches your engineer count.
| Embedded champions | Central platform team | Center of Excellence | |
|---|---|---|---|
| Fits at | Under ~15 engineers on agents | ~15 to ~150 engineers | Above ~150, or regulated at any size |
| Reporting line | Engineers stay in their product teams; no separate line | Reports to VP Eng or Head of Platform | Reports to CTO or CISO, with a written charter |
| Who approves a new prod agent | The product team's own eng lead | Platform team review | CoE sign-off against the governance charter |
| Who owns prompt-injection policy | Nobody formally; whoever the champion knows in security | Platform team lead, written down | CoE plus security, as charter policy |
| Who the CFO calls when billing doubles | Nobody owns the number (this is the shock) | Platform team owns infra cost | CoE owns the central budget and allocation |
What breaks when you hold a design too long
Embedded champions past about 20 engineers stops scaling in a way you can see in the invoice. Every team writes its own evals from scratch, so nobody can tell you whether the support agent still passes after a model swap. Security review is whoever the champion happens to know. Cost is tracked nowhere central. The billing shock lands, finance asks who signed off on $40K of tokens last month, and the honest answer is four teams each spent $10K and none of them knew about the others.
A central platform team past about 200 engineers becomes the bottleneck it was built to remove. Every new agent integration needs a ticket, the queue grows, and business units measured on shipping stand up shadow infrastructure to route around you. Now the platform team is fully booked and still blind to what runs in production.
A CoE stood up too early, under about 50 engineers, has a governance mandate and nothing to govern. Two of your strongest engineers spend most of their week on policy documents nobody reads, because the platform they are theorizing about does not exist yet and the real problem is running agents reliably. The cost is not the salaries. It is that the people who should build the platform are building slide decks.
Three signals it is time to move
Check these against your org. They are diagnostics, not recommendations.
- Embedded to platform team: three or more product teams have agents in production with no shared evals and no shared cost tracking. The second team keeps hitting incidents the first already solved.
- Platform team to CoE: a legal or compliance team has asked for a written policy on agent behavior and nobody owns it, or annual LLM spend has crossed roughly $500K with no formal approval process for adding a model or agent.
- You ran ahead of yourself: the CoE has run six months and fewer than three business units have shipped a production agent on its framework. That is governance built before the thing being governed.
When the "wrong" design is right
The bands are a default, not a law. Three cases invert them.
A regulated org in finance, healthcare, or defense may need a CoE at 30 engineers, because the compliance obligation exists at that size and does not wait for headcount to catch up. Make that CoE a governance artifact rather than a scaling hire: a one-person function plus a committee that owns the charter, not a six-person consulting org. You are buying sign-off authority; the headcount can wait.
A startup that already hired a dedicated AI infrastructure engineer at 10 engineers is running a platform team of one, whether or not anyone named it. Name it. Give that person a mandate and a reporting line so the role does not quietly revert to a floating IC every team borrows and nobody funds.
An org that decentralizes everything, with no shared platform teams for anything, should not build a central AI platform team just because AI is new. Embedded champions scale further in a decentralized org than a centralized one, because the alternative structure does not exist and will not get funded. Fighting that is a political loss you take before you write a line of code.
What breaks first is the transition, not the design
Most orgs do not fail at a design. They fail in the move between two, and both moves fail the same way: they forget to backfill.
The embedded-to-platform move breaks when the first platform hire is a senior engineer pulled off the team with the most successful agent. That team's velocity drops the week you promote them, and the person you moved spends month one feeling demoted into plumbing. Assume $200K/year fully loaded for a senior engineer and a 4 to 6 month ramp before the new structure is productive. The real cost of an unplanned transition runs $150K to $300K in lost velocity, most of it invisible because it shows up as the donor team missing a roadmap it never renegotiated. Backfill first, then promote.
The platform-to-CoE move breaks the same way one level up. The CoE gets staffed with the people who built the platform, which leaves the platform understaffed exactly when adoption is climbing and the ticket queue is longest. Name and budget the backfills before you announce the CoE, or the org chart you drew in a slide becomes a hiring gap you find in an outage.
The platform team owns your durable infrastructure: the eval harness and the CI gate that blocks a bad agent from shipping. The CoE, when it exists, owns the policy about what those gates enforce. Move the people and you move ownership of both.
What your VP of Engineering will ask
Three questions, direct answers. This is the block to forward.
"What headcount does this require?" Embedded champions is 0 net new; you are relabeling engineers you already have. A platform team is 2 to 4 to start, sized to the signal above. A CoE is 1 to 2 plus a committee if it is a governance function, or 4 to 6 if you are building a real consulting arm.
"Who owns the AI budget?" With embedded champions, product teams own their own spend, which is why the shock lands. A platform team owns infrastructure cost while product teams own usage against it. A CoE owns policy and central tooling; usage allocation back to business units varies by org, so settle it on paper before launch.
"How long does the transition take?" Embedded to platform is 2 to 4 months if the backfill happens first. Platform to CoE is 3 to 6 months. Most orgs underestimate both by 2x, because they plan the org chart before the backfills, and the org chart is the easy half.
Deciding where you sit today? Start from the invoice and the compliance calendar, not the headcount. Your band names the default design; the billing shock and the first written-policy request tell you when it stopped applying. For the full path from a funded pilot to a governed capability, the enterprise reference architecture is the layer this design plugs into.
