TL;DR — Key Takeaways

  • Enterprises often run several generations of deployment platforms at once, forcing engineers to rely on undocumented institutional knowledge to navigate them safely.
  • AI coding agents cannot depend on that tribal knowledge, turning legacy platform sprawl from an inconvenience into a potential correctness and reliability problem.
  • Platform teams should treat decommissioning as part of the definition of done, actively own difficult migrations and give legacy platforms meaningful end dates.

Somewhere in your company there is a script-based deployment setup that still runs the oldest and most critical billing services. There is a half-finished container migration from two years ago that runs most of what the company ships today. And there is the shiny new golden path announced last quarter that all new services are supposed to use.

All three are in production. All three have people on call for them. Only one of them is on the roadmap. Every one of those platforms was funded with the same promise: Standardize how software gets built and make life simpler for developers. Yet for the engineer going on call tonight, it has not gotten simpler. To debug an outage at 2 a.m., they first have to work out which era of the company a service was born in, which deployment tool belongs to that era, and which dashboards tell the truth for that particular vintage.

An engineer does that once per service and then mostly stops having to. The answer moves out of the documentation and into the team, where it survives as the sort of thing people just know. That is the quiet reason running three platforms has been survivable at all.

An AI coding agent has none of that. It reads what you actually wrote down, finds the current golden path, and applies it. Nothing in a well-maintained set of docs says “this procedure is correct unless the service predates the container migration, in which case go ask someone who was here.” The agent is not confused. It is confident, and it is working from the only version of the truth you published.

So the cost of leaving the old platform running has quietly changed character. It used to be friction, absorbed by people who adapt. It is turning into a correctness problem, and it grows with every task you hand to an agent.

The reason all three are still running is an incentive problem. Platform teams measure what they built, not what they removed.

Building the new way is some of the most rewarding work in engineering. You design a clean abstraction, leave behind the mistakes of the old system, demo a five-minute onboarding flow at the all-hands, and watch the adoption curve climb. When 70% or 80% of active services are on the new path, leadership declares victory, promotions get written, and the team moves on to the next item on the roadmap.

The remaining 20% stay behind, and that is where the real cost begins. Until the old way is actually turned off, the new platform has not removed anything. It has added a layer. Documentation now carries branches for both worlds. New hires learn the official way during onboarding and the historical way the first time they open an older repository. And the platform team quietly burns a third of its capacity keeping aging plumbing alive for a handful of holdouts.

It is easy to blame product teams for not finishing, but the last 20% are almost always the hardest, oldest, highest-risk services in the company, the ones with strange dependencies, custom networking, or original authors who left three years ago. For a product team under delivery pressure, migrating a working legacy service carries real risk and no visible reward.

It is tempting to assume agents will eventually clear that backlog. They are good at mechanical migration, and a lot of migration is mechanical. But the services still sitting on the old platform are the ones with the least documentation, the least conventional structure, and the fewest people left who understand them, which is precisely the profile that AI tooling handles worst. The help arrives in proportion to how easy the work already was.

Some of those services genuinely should not be migrated at all. A system that is frozen, contractually locked, or scheduled for retirement next year does not need a port. But there is a difference between a documented exception with an owner and an end date, and a service that has simply been left behind. The first is a decision. The second is why nobody can say when the old platform gets switched off.

Meanwhile, the other side of the ledger just got much cheaper. Scaffolding a new platform, generating the templates, writing the glue, standing up the golden path: All of it is faster than it has ever been. None of that applies to decommissioning, which is still forty conversations with teams who do not want to have them. The thing your organization was already good at got easier. The thing it was bad at did not.

That is what makes subtraction the part worth planning around:

● Count decommissioning as the definition of done. Launching a new golden path is the halfway mark. Track the share of legacy paths retired as visibly as you track features shipped, and put a number on what the old one costs: the infrastructure bill, the on-call load, the share of your team’s week spent keeping it upright.

● Own the last mile instead of assigning it. When a migration stalls on the last fifteen complicated services, more reminder emails will not move them. Sit down with those teams and do the unglamorous work alongside them.

● Give the old platform an end date and mean it. Say when support stops, who holds the pager afterward, and what happens to anything still running on it. A deprecation with no date is not a deprecation. It is just a second platform.

Anyone can add to a platform, and it has never been easier. The work that lowers the load is taking things away: Fewer ways to deploy, fewer config formats to guess between, and fewer historical layers a new engineer, or a new agent, has to know about before it is safe to let them touch anything.

Don’t pop the champagne on the afternoon the new golden path launches. Save it for the afternoon you finally turn the old one off.

Frequently Asked Questions

Why does platform sprawl create a bigger problem for AI agents?
Human engineers often learn historical exceptions through experience or colleagues. AI agents usually work from published documentation, so they may confidently apply the current golden path to services that actually require older procedures.
Why do legacy platforms remain after a new golden path launches?
The final services are often the oldest, riskiest and least documented. Product teams gain little visible reward from migrating something that already works, while platform teams are incentivized more strongly to launch new capabilities than retire old ones.
How should platform teams approach decommissioning?
Make retirement part of the success metric, quantify the cost of keeping old platforms alive, work directly with teams on difficult migrations and establish real support-ending dates for legacy systems.

SHARE THIS STORY