Skip to content
Back to the blog

One Agent per Team Is the New Multicloud

Giving every team or every task its own autonomous agent feels like freedom, but it reproduces the dynamic that made multicloud the worst practice. When building an agent is cheap, the cost moves to making the agents agree, and that cost grows much faster than the number of agents. This essay argues that the architect's lever is no longer how many agents you have, but the substrate they share: memory, identity, and a common bus to coordinate on.

Every team wants its own agent, and it's the same conversation the industry already had years ago, under a different name. Back then it was the cloud: everyone wanted their own so they wouldn't depend on the neighbor's provider, and the promise was the same one you hear now. Independence. Nobody blocking you. Moving without waiting on anyone. The private agent is sold as the end of the queue: your team stops depending on the platform team's backlog and moves at its own pace. It sounds like freedom, and that is exactly why it's hard to see that it's the same decision that made multicloud the worst practice.

The thesis is this: giving an autonomous agent to every team or every task is not free just because building the agent is now cheap. The cost didn't vanish, it moved. The cost used to be in manufacturing the agent; now it's in making all those agents agree. And that second cost doesn't grow with the number of agents: it grows much faster. The architect's lever stopped being how many agents you have and became what they coordinate on.

The drone fleet that replaces the cargo plane

Gregor Hohpe tells, in The Mighty Metaphor, the example of replacing one big cargo plane with a fleet of autonomous drones. The image is seductive: instead of one expensive, central machine with a single point of failure, you get many units that are cheap, resilient, and bought one at a time. If a drone goes down, a hundred pounds of packages go down, not the whole shipment. You scale by buying more. Each drone, on its own, is simpler than the plane.

But Hohpe uses that same example to mark where the cost hides, and that's exactly what matters here. The plane came with a pile of work we didn't see because it shipped solved from the factory: a pilot who coordinates, a single flight plan, a hold that accepts boxes of any size. When you split it into a hundred drones, that work doesn't disappear: it becomes yours and it multiplies. Now you need air-traffic control so the drones don't crash, you need to break every shipment into drone-sized packages, and you need a system that knows where each one is. The plane was expensive to buy; the fleet is expensive to coordinate.

The metaphor holds up to the test Hohpe himself demands, that the image resembling the thing is not enough, the dynamics have to match. And they match on what decides the outcome: adding one unit is cheap, but every new unit makes coordinating all the others more expensive. One more drone is a few hundred dollars; one more drone is also one more trajectory that air-traffic control has to watch against all the others. With agents it's the same. Launching a new agent today is a matter of hours; making that agent share context, identity, and results with the ones that already exist is the work that grows while nobody is watching it.

Why does the cost blow up if I only added one more agent?

Because coordination isn't paid per agent, it's paid per pair of agents that have to understand each other. And pairs grow much faster than agents.

Let's do the math, with the assumptions in plain sight so you can redo it with your own. Suppose each agent, to do its job, needs to talk to every other one: hand off context, request a result, learn what the other already did. If every pair needs its own integration channel, the number of channels is N(N−1)/2N(N-1)/2, not NN.

  • With 2 agents, 1 channel. Trivial.
  • With 4 agents, 6 channels.
  • With 10 agents, 45 channels.
  • With 20 agents, 190 channels.

You doubled the agents from 10 to 20 and the channels quadrupled. That's the trap: the agent graph is a line and the channel graph is a curve, and at the start they run almost together, so with three or four agents nobody feels the problem. It's a different cost from the one I already covered in Code Is Inventory, Not an Asset: there the cheap agent filled the warehouse of code; here it fills the map of integrations you have to coordinate. Two bills from the same drop in price.

The problem shows up when every team, reasonably, launched its own, and suddenly there are eighteen agents and a hundred and fifty-three integrations that nobody designed, nobody documented, and that break one at a time.

Now change the assumption. Instead of every agent talking to every agent, have each agent talk to a shared substrate: a place to leave and read context, a common identity, a bus the results pass through. Then each agent has a single channel, the one connecting it to the substrate, and the total is NN. With 20 agents that's 20 wires instead of 190.

Comparison of two coordination models between agents. On the left, four agents connected all-to-all, with six channels crossing each other. On the right, the same four agents each connected to a shared bus or substrate in the center, with four channels.

All-to-all, the channels are N(N−1)/2N(N-1)/2 and they explode; against a shared substrate, they are NN and grow in a straight line.

It isn't that the substrate makes the complexity disappear. It concentrates it in one place where it's designed once and maintained once, instead of spreading it across a hundred and ninety integrations that are maintained a hundred and ninety times. It's the same lesson as the lowest common denominator: multicloud didn't fail because having two providers was bad, it failed because it forced you to coordinate everything at the level of what the two had in common, and that coordination work ate the advantage of each provider. One agent per team takes you to the same place: you end up coordinating by hand, at the lowest level, what each agent could do well on its own.

I know what you're going to say

I know: every team wants its own agent so it doesn't have to wait on the one next door. If we all depend on a central platform, we're back to the bottleneck we came to remove.

It's the right objection, and part of it is correct: a badly built shared platform is a bottleneck, just as a saturated air-traffic control leaves every drone grounded. But look at what you're comparing. The shared substrate is not a central team that approves every action; it's a set of primitives on top of which each agent moves by itself. Air-traffic control doesn't fly the drones: it gives them a common space, precedence rules, and a way to know where everyone is, and then each drone flies. Autonomy isn't touched; what's shared is the ground it's exercised on.

And there's an asymmetry the rush hides. Skipping the substrate accelerates today: your team launches its agent this week without asking anyone for anything. But the channel you didn't design doesn't disappear, it's inherited by whoever comes later to make your agent talk to theirs, and they pay it with interest. It's the same dynamic as autonomy as a blank check: freedom is granted cheap and clawed back expensive. A loose agent ships in an afternoon; untangling eighteen agents that integrated pair by pair, each its own way, is a quarter-long project.

And before you think it: no, the answer is not to forbid teams from having agents and centralize everything into a single one. That's going back to the cargo plane, with its single point of failure and its waiting queue. The drone fleet is a good idea; what decides whether it works is that air-traffic control exists before there are a hundred drones in the air, not after the first crash.

Two panels labeled as different options, «ONE AGENT PER TEAM» and «ONE SHARED PLATFORM», each with an arrow leading down to the same identical box: «N agents, N(N-1)/2 channels, someone has to coordinate them».

The dilemma is false: the question isn't how many agents, but what they coordinate on.

Is this an industry pattern or a belief of mine?

It's a pattern, and the best evidence is that the industry itself is already building the substrate before most people notice they need it. Amazon Bedrock AgentCore doesn't sell «an agent»: it sells the primitives many of them run on: a common identity so every agent authenticates the same way, a memory and a context that are shared instead of rebuilt in each agent, and a gateway the tools come in through. It is, literally, air-traffic control for drone fleets. The product isn't the drone; it's the airspace.

The same thing shows up from the other end with Kiro and its spec-driven development: the spec is the common ground several agents read so they don't contradict each other, the document they coordinate on without having to talk to one another. In both cases the value stopped being in the individual agent and moved to the substrate they share.

And there's a reason older than all of this why the problem shows up even if nobody designed it. If every team builds its own agent, the agent system ends up copying the org chart: as many agents as teams, integrated as badly as the teams talk to each other. It's Conway's law, which Melvin Conway stated in 1968: a system ends up with the shape of the communication structure of whoever builds it. One agent per team isn't an architecture decision, it's the org chart leaking into the software. And the org chart was never a good architecture diagram.

One agent per team doesn't give you freedom: it gives you multicloud's coordination tax, spread across integrations nobody signed off on.

My suggestion

On Monday, bring a single question to your next architecture review: before approving the next agent, what shared substrate is it going to coordinate on with the ones we already have? If the answer is «it integrates straight into the data team's agent», you just authorized one more channel on the curve, not one wire to the bus. It's not that the new agent is wrong; it's that you're hanging it in the wrong place.

And if you want something even more concrete, do the exercise with numbers in the next meeting: count how many agents are in production today, count how many point-to-point integrations connect them, and project both curves out six months at the rate teams are launching agents. If the integration curve is already steeper than the agent curve, you don't have a speed problem: you have a warehouse of channels, and the shared substrate is the decision that flattens it. Decide it before drone number one hundred, not after the first crash. ;)

References

  1. The Mighty Metaphor — Gregor Hohpe
  2. Amazon Bedrock AgentCore — AWS
  3. Kiro — spec-driven development
  4. Conway's Law — Melvin Conway
  5. Conway's Law — Martin Fowler
  6. Multicloud, the Worst Practice
  7. Autonomy Is a Blank Check
  8. Code Is Inventory, Not an Asset