Beer & Servers Don't Mix

Learning Scale from Constraints

Twenty years ago, before AWS existed, before anyone had heard the word “cloud,” there was a simple rule in tech: You bought iron or you died. Period. No elastic scaling. No pay-as-you-go. You guessed what you’d need, you wrote a check, and you prayed it’d last 5 years so you got moneys worth.

I’m sitting in a conference room at an ecommerce startup in Australia, and I’ve just finished what I think is the presentation of my career. Half a million dollars of blade chassis gleaming on the PowerPoint behind me. State-of-the-art switches. A database server that could handle anything we threw at it for the next five years. My dev team is nodding. The accounts guy is taking notes. We’re about to transform this company.

Then one of the founders says, “Can we have the room?”

Here’s what you learn about startup founders: They’ve already done the math before you walk in. They know exactly how much runway they have. Down to the day.

The founder — let’s call him Dave — leans forward. He’s got that look. The one that says he’s about to explain reality to someone who’s been living in a fantasy.

“Profits haven’t been what we projected,” he says. Translation: We’re bleeding cash. “The peak sales season” — two months of summer sales — “that’s what gets us through the year.” Translation: Miss this, we’re dead. “If we can’t get through this with 100k, we’re in trouble. Big trouble.” Translation: Your beautiful half-million-dollar plan just became a $100,000 problem.

Then he says the thing he doesn’t want to say: “If we can’t make this work, we’ll need to make some difficult decisions about headcount.”

I’m doing the math in real-time. My team is four people including me. At $100K, we can’t even buy two of the blade chassis I spec’d. We definitely can’t buy the database server. We can’t buy anything that matters.

But here’s the thing about constraints — sometimes they force you to see what’s been there all along.

I’m lying in bed that night, and I’m thinking: What if we’re solving the wrong problem? What if the question isn’t “How do we build infrastructure for five years?” What if it’s “How do we survive two months?”

Two months. Eight weeks. Sixty days.

That’s when I discover eBay has a section for decommissioned government servers.

Let me explain what buying enterprise hardware on eBay meant in 2004: It meant you were either desperate or insane. These servers came with no warranty. No support contract. No guarantee they’d even boot up. One seller’s listing said something along the lines of: “Sold as-is from government surplus. Buyer responsible for wiping drives.”

“We’re really doing this?” Bruce, our infrastructure guy, asks me.

“What’s our alternative?”

He loads up the eBay website in Firefox.

The servers arrive in the back of a white van. Bruce calls me from the pickup location — some guy’s storage unit in an industrial park. “They’re wrapped up in bubble wrap like… like big sausage rolls,” he says.

“Do they look… functional?”

“They look like servers that have seen things.”

We’re paying $8,000 for what would have cost us $80,000 new. If even half of them work, we’re ahead.

Here’s what Dell, IBM or HP won’t tell you about their blade chassis: They’re overengineered. They’re built to run in defense contractors’ data centers for a decade. They’re built to survive. Which means when the Department of Defense sells them after four years, they’ve got another four years in them. Maybe six. You just have to be willing to bet your company on it.

One week before the summer sale season kicks off, a blade fails. Just dies. No warning, no gradual decline. Dead.

Bruce is on the phone: “What do you want me to do?”

And I hear myself say something that would have been heresy six weeks earlier: “Pull it. Leave it dead. We’ve got thirty-five more.”

This is the moment. Right here. This is when we discover web scale.

Because here’s what nobody understood yet: You don’t need perfect hardware. You need expendable hardware. You need hardware you can afford to lose.

The first weekend of summer sales, we’re watching the metrics. Traffic is up spiking as high as 50x in some minute windows. The janky eBay servers are handling it. They’re not just handling it — they’re cruising. We’ve got easy 20% capacity to spare during spikes.

Monday morning, Dave pulls me aside. “How much did we spend total?”

“One-twenty. Could have done it for one-ten, but we bought some insurance.”

“Insurance?”

“Extra garbage servers from eBay.”

He’s looking at me like I just explained how we turned lead into gold.

The next year, same problem, but now we know the game. Bruce calls the data center: “You got any customers coming off lease? Any hardware they’re abandoning?”

Turns out, every data center in 2005 is sitting on piles of abandoned hardware. Four-year-old servers that companies walked away from because the new models were 30% faster. We could lease them for almost nothing.

Year two: We lose four blades the week before peak season. We don’t even blink. Pull them, redistribute the load, keep going. The application layer gets smarter because it has to. The hardware becomes commodity because we treat it that way.

By year three, we’re running massive traffic on what is essentially technological garbage. Our competitors are signing million-dollar Dell contracts. We’re bidding on government auctions.

Here’s what I didn’t know in that conference room, staring at my beautiful, doomed PowerPoint: We weren’t just solving a budget crisis. We were stumbling into the future. Every hyperscale company today — Google, Facebook, Amazon — they all learned the same lesson we learned in that storage unit, unwrapping those sketchy servers:

The hardware doesn’t matter. The hardware is already dead. Plan for it. Design for it. Embrace it.

Scale isn’t about buying the best servers. It’s about buying so many terrible servers that it doesn’t matter when they fail.

Here’s what nobody tells you about running hundreds of servers: Failure isn’t a possibility — it’s a mathematical certainty. At ten servers, you might lose a disk every few months. At a hundred, you’re losing disks weekly. At a thousand? Daily. Multiple times daily. The bigger you go, the worse it gets. It’s not exponential, it’s relentless.

So you stop thinking about preventing failure. You stop caring about individual disks dying. Then you stop caring about entire servers dying. You just assume they’re already dead and design your application accordingly.

And here’s the kicker — the software changes you need? They’re not huge. A load balancer here. Some health checks there. Replicate your data. Retry your requests. Suddenly your application doesn’t care that Server 47 just burst into flames. It routes around damage like the internet was supposed to. Much of this is now built into modern public cloud platforms, and its these exercises of discovery was the basis for it.

We discovered this in a storage unit with sketchy eBay servers because we had no choice. Google discovered it with thousands of cheap commodity boxes. Facebook discovered it when they realized enterprise hardware couldn’t scale to a billion users.

Same lesson, different budgets: Once you accept that everything will fail, you’re finally free to build something that won’t.