Beer & Servers Don't Mix

The Ownership Illusion: Why Your Platform Team Is Building Solutions to Problems That Don’t Exist

Or: How We Spent Three Years Maintaining a Band-Aid on a Bullet Wound

It was 2:47 PM on a Tuesday, and the queue had sixty pull requests in it. Sixty. The Slack channel was a wall of red notifications, and across the meeting room table, a dozen tech leads sat with arms crossed and faces that could curdle milk. One of them had brought a printed spreadsheet — actual paper — showing how many story points his team had lost that week. I remember thinking the only thing missing from this scene was someone flipping a table.

This wasn’t supposed to happen. We’d just launched our revolutionary new merge system, the one we’d spent months building, the one that was going to save everyone. Instead, nothing had merged for an entire day. The platform team sat hunched over laptops, debugging furiously. The product owners were in their standups learning that their carefully planned sprints had just been torpedoed. And I was starting to understand something uncomfortable about how we solve problems in engineering organisations.

We’d built something brilliant. And it was exactly the wrong solution.

The Merge Problem Nobody Asked Us to Solve

Let me take you back to where this started. We had a large monorepo — a single repository with ownership divided among fifteen teams. Hundreds of engineers, one codebase, one master branch. You can probably already see where this is going.

The problem was deceptively simple. Two engineers try to merge to master at roughly the same time. The first one’s CI passes, they merge successfully. The second one now has to cancel their CI pipeline, pull the new changes from master into their branch, rerun CI, and then try again. If someone else merged while they were doing that, the cycle repeats.

With fifteen teams and hundreds of engineers, this synchronous process created a queue. On an average day, twenty to thirty people would be waiting. We’d merge about once every forty-five minutes, running twenty-four hours a day. Sunday was the only time the queue ever emptied completely.

The whole thing was orchestrated through a Slack bot. Engineers would type commands to join the queue, check their position, or bail out when they’d been waiting too long. It was clunky, frustrating, and clearly unsustainable as we continued to grow.

Something needed to be done. Everyone agreed on that.

Enter the Platform Team

So we did what modern engineering organisations do when they face a cross-cutting technical challenge: we formed a platform team. Their mission was clear — solve the merge problem and unblock our engineers.

Here’s the critical detail that nobody thought about at the time: this team didn’t own any parts of the repo. They weren’t empowered to make changes to the existing code or system. They were positioned outside the problem, tasked with observing it from a distance and building something to fix it.

And so, exactly as Conway’s Law predicts, they did what any team in their position would do. They built a new system on top to “manage” the disaster.

The computer scientist Melvin Conway observed in 1967 that organisations are constrained to produce designs which mirror their own communication structures. We had created a team that couldn’t touch the underlying system, so naturally they built a layer that sat above it. They solved the problem they were empowered to solve, not the problem that actually needed solving.

What we built was, in essence, what GitLab now offers as their merge train feature — parallel CI runs on stacked branches, automatically managing the queue so that multiple pipelines could validate simultaneously. Ours was different in some ways, a bit better in others, and built on top of GitHub rather than GitLab.

By parallelising CI on top of stacked branches, we increased merge frequency significantly. A feat of amazing engineering. A shiny new tool from our shiny new platform team.

Wrong.

The Illusion of Progress

As the economist John Kenneth Galbraith once observed, “The conventional view serves to protect us from the painful job of thinking.” We had convinced ourselves that our conventional view — the merge queue is the problem, optimise the merge queue — was the right one. It protected us from the painful reality that we were optimising the wrong thing entirely.

We hadn’t fixed the underlying problem. The single point of contention — the monorepo itself — remained unchanged. We’d just widened the bottleneck through brute force. Like responding to a clogged drain by installing a bigger pump rather than clearing the blockage, we’d thrown engineering horsepower at a problem that didn’t need horsepower — it needed a plumber.

In 2021, researchers at the University of Virginia published a fascinating study in the journal Nature. They gave participants a simple task: stabilise a Lego structure. The structure had an uneven roof, and participants could either add bricks or remove them. Most people instinctively added bricks to prop up the weaker side. Very few thought to simply remove the single block causing the instability in the first place.

Across eight different experiments, the researchers found that people systematically default to searching for additive solutions and overlook subtractive ones. As the researcher Benjamin Converse explained, “Additive ideas come to mind quickly and easily, but subtractive ideas require more cognitive effort. Because people are often moving fast and working with the first ideas that come to mind, they end up accepting additive solutions without considering subtraction at all.”

This is what we did. We added a system. We added complexity. We added a team to maintain that complexity. We never seriously considered removing the thing that caused the problem in the first place.

The Launch That Launched a Thousand Complaints

That Tuesday meeting room, with the crossed arms and the printed spreadsheets and the barely contained fury — that was the second day of our new system. The whole first day, nothing had merged. The queue had ballooned to sixty pull requests. Deployments were blocked. Product work was delayed. The POs had learned about it in standup, and their engineers had carried that frustration directly into the meeting.

I was genuinely surprised they didn’t bring knives.

Here’s the thing that made it so much worse: it was a solution they weren’t brought into. It was a system pushed down from management to save them. And when it did the opposite on day one, this is what happens.

We spent a week getting it operational, another two weeks getting on top of the queues, and then three years maintaining it. Three years of a dedicated team keeping this elaborate machinery running. Three years of workarounds, edge cases, and increasingly complex logic to handle scenarios the original design hadn’t anticipated.

The Solution That Actually Worked

Years later, we finally did the thing we should have done from the start. We broke the monorepo apart. We split it into twenty-two pieces, each given to a separate team.

The result? Most of these pieces max out at five pull requests a day. Five. They don’t need a separate team to build tooling for them. They don’t need a merge train. They don’t need a queue at all.

We share code that infrequently changes via libraries. We share code that frequently changes via APIs. It works.

The transformation wasn’t just technical — it was organisational. Each team now owns their piece end-to-end. They can merge when they want, deploy when they want, and make decisions without coordinating with fourteen other teams. The bottleneck didn’t just get faster; it evaporated.

Looking back, the irony is painful. We spent three years maintaining an elaborate system to optimise merging into a single repository, when the actual solution was to stop having a single repository. We built a faster Band-Aid when we needed surgery.

What We Got Wrong About Ownership

Beyond the psychology of addition over subtraction, we made another critical error: we didn’t ask the people closest to the problem.

The teams that owned the repo, that worked in it every day, that felt the pain of the merge queue in their bones — we didn’t bring them into the solution. Instead, leadership put ourselves in a room, came out with the solution like Moses with stone tablets, and handed it to a platform team with a mandate to execute.

In retrospective, what I would have done — and what we eventually did years later — was take people from those teams, ask them what they thought we should do, and then empower them to fix the problem themselves.

The author and management consultant Stephen Covey put it well: “What you do has far greater impact than what you say.” We said we valued our engineers’ expertise. What we did was exclude them from the most important decision affecting their daily work.

The Platform Team Paradox

None of this is to say platform teams are inherently bad. They’re not. But there’s a particular failure mode that’s worth understanding.

A platform team without ownership of the systems they’re trying to improve is structurally incentivised to build layers on top rather than fix underlying problems. They add observability. They add automation. They add abstraction. Each addition makes perfect sense in isolation, each represents genuine engineering effort, and collectively they can create a thicket of complexity that obscures rather than resolves the original issue.

The golfer Jack Nicklaus once said, “Learn the fundamentals of the game and stick to them. Band-Aid remedies never last.” We were applying Band-Aid after Band-Aid, each one a little more sophisticated than the last, when the fundamental problem remained untouched.

If your platform team is building tools to help other teams cope with a problem, ask yourself: why are we coping? What would it take to eliminate the need to cope entirely? And crucially: who actually has the power to make that change?

The Questions That Might Have Saved Us

Looking back, there are questions we should have asked before forming that platform team:

Is the problem we’re solving actually the root cause, or is it a symptom of something deeper?

Do the people we’re empowering to solve this problem have the authority to address the underlying system, or only to build on top of it?

Have we asked the engineers who live with this problem daily what they think the solution should be?

If we magically made this problem go away tomorrow, what would be different? Is there a way to just… do that?

Are we adding because adding feels productive, or because adding is actually the right answer?

I wish I could tell you we’d learned these lessons easily. We didn’t. It took three years and a lot of wasted effort before we finally stopped optimising the merge queue and started questioning why we had a merge queue at all.

The Uncomfortable Truth

Platform teams are supposed to accelerate product teams. But when they’re positioned outside the systems they’re meant to improve, they can become something else entirely: sophisticated caretakers of problems that should have been eliminated rather than managed.

Conway’s Law tells us that systems mirror communication structures. When you create a team that can only communicate with a problem through abstraction layers, don’t be surprised when they build more abstraction layers.

The Lego researchers found that when participants were explicitly told “removing pieces is free,” far more of them chose the simpler subtractive solution. Sometimes the most important thing leadership can do is explicitly authorise removal, explicitly empower simplification, and explicitly ask: what if we just… didn’t have this problem?

Because the fastest merge queue is the one you don’t need at all.

Now, if you’ll excuse me, I need to go explain to a platform team why their latest proposal for improving our deployment pipeline might benefit from talking to the teams who actually deploy things. Old habits, it turns out, die remarkably hard.