The Lego Problem: Why Engineers Add Complexity Instead of Removing It
Or: Why Your Architecture Reviews Have a Benefit Slide and No Cost Slide
It was 9:52 on a Wednesday morning and someone had drawn a new box on the whiteboard. The meeting room on level 7 still carried the stale warmth of the 10 AM standup that had run long, and the aircon hadn’t caught up yet. The box had a name — something with “orchestrator” in it — and three arrows pointing to existing services. The engineer presenting it was animated, enthusiastic, already two slides into the “how.” Nobody had asked “whether.” The product owner was nodding. The tech lead was sketching integration points on a napkin. And somewhere in the back of the room, a senior engineer was staring at the whiteboard with the quiet resignation of someone who knew they’d be on-call for this thing by March.
That box was never questioned. It was added, deployed, monitored, maintained, and is still running today. And the problem it was supposed to solve? It’s still there too — just with an extra component sitting on top of it.
This isn’t a post about microservices being bad. It’s not a post about simplicity as some kind of spiritual practice. It’s about a cognitive bias that’s been measured, replicated, and published in Nature — and that nobody in our industry talks about, despite it explaining half the architectural decisions that keep us up at night.
The Experiment That Explains Your Architecture
In 2021, researchers at the University of Virginia ran eight experiments with over 1,500 participants and found something that should be required reading for every engineering leader. When asked to improve something — anything — people systematically default to adding. Not because they considered subtraction and rejected it. They literally didn’t think of it.
The most striking experiment involved Lego. Participants were given a structure with a wobbly roof supported by a single off-centre pillar. The task: make it stable enough to hold a brick on top. Each added block cost 10 cents. Removing the pillar was free — the roof would rest flat on the base, perfectly stable, zero cost. But most people reached for more bricks.
As Antoine de Saint-Exupéry wrote in Wind, Sand and Stars, “Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.” He was writing about aircraft design — another field where unnecessary complexity kills. And he published that in 1939, which means we’ve had nearly a century to absorb this lesson and we’re still adding pillars.
Now apply that to your last architecture review. Someone proposed a new service, a caching layer, a message queue. The discussion was about how to add it — API contracts, team ownership, deployment strategy. Subtraction never made it to the whiteboard. Not because anyone argued against it. Because nobody’s brain put it there.
The Additive Default in Software
You’ve sat in these meetings. We all have. The pattern is so consistent it’s almost formulaic.
“Performance is slow” → add a caching layer. “Services are coupled” → add a message queue. “Deployments are risky” → add a service mesh. “Auth is inconsistent” → add an API gateway.
Each addition sounds reasonable in isolation. Each one passes code review. Each one ships. And each one adds a component that can fail, an integration that needs monitoring, a system that needs maintaining — forever.
Here’s what gets discussed versus what doesn’t:
When you propose a new service, you talk about the API contract and team ownership. You don’t talk about the deployment pipeline, monitoring, on-call rotation, documentation, and inter-service testing that come with it. When you propose a caching layer, you talk about the performance improvement. You don’t talk about cache invalidation logic, consistency bugs, memory management, and cold-start problems. When you propose a message queue for decoupling, you don’t talk about ordering guarantees, dead letter handling, replay logic, and lag monitoring.
The pattern: every proposal has a benefit slide and no cost slide. Not because the proposer is lazy, but because the costs are invisible at proposal time and only become real at 3 AM six months later.
I wrote about the long-term decay this causes in Code Entropy. But entropy is the effect. Additive bias is the cause. Your brain is wired to make your codebase more complex, and it’s doing it without your permission.
The Compound Cost Nobody Calculates
This is where it gets worse than “more stuff = more problems.” Complexity compounds.
Two services with a message queue between them isn’t three components — it’s an exponential increase in failure modes. Service A can fail. Service B can fail. The queue can fail. A can fail while B is fine. B can fail while A is fine. The queue can lose messages. The queue can deliver duplicates. The queue can deliver out of order. Messages can poison the queue. A can produce faster than B consumes. And all of these interact with each other.
Architecture Components Integration Points Potential Failure Modes Monolith 1 0 Low 3 services 3 3+ Medium 10 services 10 20+ High 50 services 50 100+ Very High
Here’s the uncomfortable truth that most microservices advocates gloss over: microservices don’t reduce complexity. They redistribute it. From code complexity to operational complexity. From compile-time problems to runtime problems. From problems you find in development to problems you find in production. From problems your IDE catches to problems your monitoring dashboards may or may not surface, depending on whether you’re measuring the right things.
The legendary basketball coach John Wooden once said, “Don’t mistake activity for achievement.” We’ve been remarkably active in our architecture reviews — adding services, adding layers, adding abstractions. The question is whether any of it achieved the simplicity we actually need to move fast and stay reliable.
Kent Beck and the Forgotten Fourth Rule
Most engineers know Kent Beck’s four rules of simple design. Or rather, they know three of them:
- Passes all tests- Reveals intention- No duplication- Fewest elements That last rule — fewest elements — is the subtraction rule. It’s not “fewest useful elements” or “fewest elements given our current architecture.” It’s fewest. Period. If you can remove something without violating rules 1–3, you should.
This connects to Dan McKinley’s influential essay Choose Boring Technology — the argument that every technology choice has an “innovation token” cost, and most organisations only have a few tokens to spend. Each addition to your stack spends a token whether you acknowledge it or not. That message queue you added? Token spent. That caching layer? Another token. That custom auth service? You’re out of tokens and you haven’t even started building the thing your users actually want.
The Distributed Monolith: Adding Everything, Gaining Nothing
The worst outcome of additive bias in microservices is something every architect has seen and nobody wants to admit they’ve built: the distributed monolith. You’ve added all the operational complexity of distributed systems while gaining none of the benefits. Services that must deploy together. Shared databases. Synchronous calls everywhere. No independent deployability.
You got the cost of microservices with the coupling of a monolith. This is what happens when the answer to every problem was “add a service” and the answer to no problem was “merge two services back together.”
Segment learned this the hard way. They grew to over 140 microservices, each with its own testing, deployment, and monitoring overhead. Three full-time engineers were spending most of their time just keeping the system alive instead of building features. They consolidated back into a monolith and saw immediate improvements in developer productivity — going from 32 improvements to shared libraries in a year to 46 improvements the year after consolidation. As their engineer Alexandra Noonan put it at QCon London, incorrectly implemented microservices can leave you unable to do product development because you’re drowning in the complexity.
This isn’t an anti-microservices argument. Microservices work if you do them right, and when you have a large enough scale of engineers, they’re genuinely needed. But the decision to add a service should require the same scrutiny as any other architectural decision — and that means asking the question we almost never ask.
The Removal That Actually Helped
Let me give you a real example. When we started building the new micro frontends for the YCS monolith split at Agoda, I made one decision early: no session state. All APIs would be stateless until someone could show me a concrete reason we needed it.
People were sceptical. Session state was just how things were done. The assumption was baked in so deeply that questioning it felt almost rude. But we held the line.
We spent two years migrating endpoints out. And nobody cared. It worked. We didn’t need state. We had a couple of cookies and that was it. Two years, hundreds of endpoints, and the subtraction of session state — something we’d been carrying around like a security blanket — caused zero problems while eliminating an entire category of complexity.
That’s what subtraction looks like in practice. Not dramatic. Not heroic. Just the quiet absence of a problem that never needed to exist.
And here’s the kicker: after two years, someone finally came to me with a legitimate use case for session state. We added it. And because we understood exactly why we were adding it and what problem it solved, we were able to drop our P99 server-side response time SLOs by whole seconds across multiple services. Whole seconds. Not milliseconds — seconds. That’s the difference between adding something because “we might need it” and adding something because you’ve proven you do. When you need it, you add it. Not before. YAGNI isn’t just a catchy acronym — it’s a strategy that pays compound interest when you finally do spend.
Patch on Top of Patch
Here’s the flip side — what happens when additive bias runs unchecked.
I was in a standup once where we were discussing a flow for onboarding new clients. There was a bug somewhere in the pipeline — about one in a hundred clients would end up with incorrect field data. The product owner asked if we could write a script to fix the bad data. Fair enough. We did — a simple SQL script, took no time at all.
Then someone said, “Now let’s run it on a schedule.”
I need you to sit with that for a moment. The proposed solution to a bug that corrupted one percent of client data was not to find and fix the bug. It was to run a script on a schedule that would periodically clean up after the bug. Patch on top of patch. A scheduled job to compensate for a defect that nobody wanted to dig into.
This is the stereotypical pattern I see everywhere. Rather than fixing root causes, we add another layer — a script, a service, a workaround, a retry mechanism, a compensating transaction. And then your system ends up as a tower of patches, each one depending on the others, none of them addressing the original problem. I’ve written about this pattern before — how temporary things become permanent when nobody asks the next question.
E.F. Schumacher, the economist, wrote in Small Is Beautiful: “Any intelligent fool can make things bigger, more complex, and more violent. It takes a touch of genius — and a lot of courage — to move in the opposite direction.” That quote lands because it names the emotional truth: subtraction requires courage. Proposing to remove something means challenging a previous decision, possibly your own. It means saying “we got this wrong” or “this isn’t needed anymore.” That’s harder than proposing something new. And that’s exactly why it’s the more valuable skill.
The Questions to Ask Before Adding Anything
From handling large-scale failures over the years, I’ve learnt to love simplicity at scale. I don’t want smart things. I want dumb things that work reliably when everything else is on fire. The best design pattern at scale is simplicity. Not because it’s aesthetically pleasing — because it’s the only thing that survives contact with production at 3 AM.
I’ve been on the other side of this. Back in my previous life in event ticketing, when a major event went on sale at 9 AM, traffic would spike 100x in thirty seconds. Not gradually — thirty seconds. AWS autoscaling can’t respond that fast. I’ve watched it try. Your clever distributed caching strategy, your elegant circuit breakers, your beautifully abstracted service mesh — none of it matters when a hundred thousand people hit your system simultaneously and the infrastructure hasn’t caught up. The dumb things survived. The smart things fell over. As the military saying goes, no plan survives contact with the enemy — and neither does your architecture when the load arrives faster than your scaling policy.
Here are the questions that should be mandatory in every architecture review:
“Can we solve this by removing something instead?” — Start here. Always. Your brain won’t suggest it naturally, so you need to force it onto the whiteboard.
“What’s the simplest solution that works?” — Not the most elegant. Not the most scalable. Not the most “proper.” The question isn’t “how can we do this faster” — it’s “how can we do this in a simpler way.”
“What’s the cost of this addition over 5 years?” — Who maintains it? Who’s on-call? What happens when the person who built it leaves? Every component you add is a commitment your future team is making on your behalf.
“What would happen if we did nothing?” — Sometimes the answer is “nothing bad.” That’s not laziness — that’s a signal.
“Would I be comfortable explaining this in a post-mortem?” — If the answer is “we added it because it was interesting,” you have your answer.
The most valuable architectural skill isn’t knowing what to build — it’s knowing what to remove. And that skill is rare not because it’s difficult, but because it’s cognitively unnatural. You have to actively fight your own wiring to even consider it.
The Bottom Line
We quote “keep it simple” like a mantra, but we don’t live it. Sophistication in engineering culture means more patterns, more abstractions, more services. The engineer who proposes removing something is seen as less sophisticated than the one who proposes adding a new distributed transaction coordinator. We have it exactly backwards.
Your brain is working against you. The research is clear — you will default to adding, not removing, and you won’t even notice you’re doing it. The only defence is making subtraction a deliberate, first-class part of every architectural conversation. Ask the removal question first. Celebrate deleted code. Treat merging services back together as a sign of maturity, not defeat.
Microservices, caching layers, message queues, orchestrators — they’re all tools, and sometimes they’re the right tools. But every one of them should have to justify its existence against the alternative of doing less. And “doing less” should always be on the whiteboard.
Now, if you’ll excuse me, I need to go review an architecture proposal that adds three new services to solve a problem that might also be solved by deleting one. I already know which option requires more courage. I’m working on it.