Beer & Servers Don't Mix

The Onboarding Cliff: Why Your New Engineers Are Productive at Month Six (If They’re Still Here)

Or: What Restaurant Kitchens Mastered That Tech Companies Haven’t

Somchai walked into Monday’s standup with the easy confidence of someone who’d estimated a four-hour task. Small change to a backend API — add a field, update the mapping, ship it. Coffee in hand, headphones around his neck, he told the team he’d have it done by end of day. The Jira card was already in the “In Progress” column before he’d sat back down.

By Tuesday’s standup, the confidence was gone. He’d spent the entire previous day trying to get the backend API running on his laptop. There was a Confluence page — long, dense, last updated by someone whose Slack status now read “alumni” — that listed the steps to make it work locally. He’d followed them. It didn’t work. He’d followed them again, more carefully. Still nothing.

By Wednesday and Thursday, he was still at it, and now other people were burning time too. Engineers from the team that owned the service were huddling around his desk, trying to help. They tweaked JVM settings. They tried building inside a Docker container. They started commenting out parts of the sbt file to isolate the problem. The PO was asking questions in standup. His manager pulled him aside to check on progress. Everyone was being nice about it, but the message was clear: this is taking too long.

By Thursday afternoon, Somchai gave up on running it locally. He started making his code changes blind — pushing to Git, waiting twenty to forty minutes for the CI pipeline to tell him whether it worked, adjusting, pushing again. Each iteration cost him half an hour of staring at a progress bar, and the frustration was visible from across the floor.

Day five. Standup. Somchai had taken four days to complete what he’d estimated — and what genuinely was — four hours of actual work. He was tired, frustrated, and still waiting for a code review from the other team. He did not, to put it mildly, have a good developer experience.

Here’s the thing: Somchai wasn’t new. He wasn’t a graduate. He was a competent engineer who’d simply never needed to work in that particular service before. And the moment he stepped outside the code he already knew, he hit a wall that nobody had bothered to fix because everyone who worked there daily had long since learned to climb over it.

Your onboarding isn’t broken because your documentation is bad. Your onboarding is broken because your systems are bad, and onboarding is just the moment that makes it impossible to pretend otherwise.

The Litmus Test Nobody Wants to Take

Sometime in 2009, in a meeting room that smelled like instant coffee and fresh whiteboard markers, a consultant from Readify — a well-regarded Australian software consultancy — was leaning forward in his chair, gesturing at a diagram he’d drawn in blue marker on the whiteboard. He had that energy, the kind where someone is so convinced of what they’re telling you that they can’t sit still. “End of day one,” he said, tapping the whiteboard. “Clone, open, F5. They should be pushing code to production by the afternoon.”

His colleague was nodding along, barely containing the same enthusiasm. I looked at my colleague. My colleague looked at me. Our reality at the time involved a Visual Studio solution that took eleven minutes to build, a setup guide we’d printed and stapled — because we didn’t trust the wiki — and a new starter ritual that required at least two days of someone sitting next to you explaining which projects in the solution you could safely ignore. TFS was our source control. We were still arguing about whether ReSharper was worth the RAM.

The gap between what those consultants were describing and what we were living felt like science fiction. But the confidence in that room planted something — a kaizen goal I’ve been chasing for close to twenty years now.

Can a new joiner push code to production by end of day one? That question sounds absurd to most engineering leaders the first time they hear it. It certainly sounded absurd to me. But they showed me examples where it worked — environments with well-designed services and clear business domain separation, where each piece of the system was simple enough that someone could onboard, understand it, and contribute meaningfully within a single day.

That’s not a fantasy. That’s a design goal. And the distance between where you are and that goal is a precise measurement of the hidden complexity your existing engineers are silently coping with every single day. Just like Somchai was.

As Michael D. Watkins wrote in The First 90 Days, “The first 90 days are not about proving yourself. They’re about learning how to be effective in a new environment.” We’ve somehow inverted this entirely — we treat onboarding as a test of the new hire’s capability rather than a test of our system’s comprehensibility. When someone struggles for three months, we question their talent. We should be questioning our architecture.

The Kitchen That Already Solved This

There’s a reason the subtitle of this post references restaurant kitchens, and it’s not just because I enjoy a good food analogy.

The French brigade system, developed by Auguste Escoffier in the late 1800s, solved the onboarding problem over a century ago. A new line cook walks into a professional kitchen and can be productive on their first service. Not because they’re a culinary genius — because the system is designed for it. The station is standardised. The mise en place — everything in its place — means every ingredient, every tool, every component is visible, labelled, and within reach. The feedback loop is immediate: you plate a dish, the chef sees it, you know within seconds whether it’s right.

Compare that to most engineering organisations, where the “kitchen” is a sprawling landscape of services with no map, the “ingredients” are scattered across repositories nobody owns, and the feedback loop is “run the pipeline and come back in forty minutes to find out it failed on a linting rule nobody told you about.”

The chef and television personality Anthony Bourdain, who spent decades in professional kitchens before becoming famous, had a phrase he used constantly: mise en place. Everything in its place. He treated it almost as a philosophy of life, not just cooking. The question for us is: what does mise en place look like for a software engineering team? What would it mean for every tool, every configuration, every piece of context to be exactly where a new person expects to find it?

Microservices Win This One (And It’s Not Even Close)

I’ve heard the argument that microservices make onboarding harder — more services to understand, more moving parts, more complexity. In my experience, it’s precisely the opposite.

When your new hire sits down to troubleshoot a problem in a monolith, they’re staring at thousands of lines of NX config before they’ve even reached the business logic. The build system alone is an archaeological dig. And we haven’t even mentioned the CI pipeline yet — that’s another layer of complexity that exists purely because the monolith demands it.

With well-designed microservices, a single service might have a hundred lines of Vite config. The domain is bounded. The scope is comprehensible. A new engineer can hold the entire context in their head, which is precisely the point. You’re not asking them to understand everything on day one — you’re asking them to understand one thing well enough to contribute. That’s a fundamentally different proposition.

The key phrase there is “well-designed.” Microservices that are just a distributed monolith with extra network calls won’t help. But microservices with genuine business domain separation, clean boundaries, and simple tooling? That’s where the onboarding magic happens.

The Intern Test

Here’s where I get to share something I’m genuinely proud of, even though it’s still a work in progress.

Every summer, we take on twenty to thirty interns for an eight-week programme. Two weeks of training, then three two-week sprints. We form them into teams, and by sprint one — that’s week three — they need to be producing production code. These are students, typically with a reasonable amount of React knowledge and a touch of Java, so getting them to work on C# and our React front end isn’t an unreasonable stretch. But they’re not senior engineers. They don’t have years of pattern recognition to fall back on. They can’t power through poorly designed systems on sheer experience.

After a few years of iterating on the content, we got the onboarding down to one week last year.

Now, when I hire a senior engineer, onboarding is naturally faster — they bring experience, they can navigate complexity, they pattern-match against systems they’ve seen before. But that’s exactly why I don’t benchmark myself on senior hires. I benchmark on the interns. Because interns reveal what seniors silently absorb.

If an intern struggles with something, that’s not the intern’s failing. That’s hidden cognitive load — complexity that exists in your system that your regular engineers have learned to cope with, but shouldn’t have to. The intern can’t compensate. They don’t have the muscle memory. So when they hit a wall, you’ve found a wall that everyone is hitting; the experienced engineers have just learned to climb it so reflexively they’ve forgotten it’s there.

As the physicist Richard Feynman once said, “The first principle is that you must not fool yourself — and you are the easiest person to fool.” We fool ourselves every day about how simple our systems are, because we’ve adapted to their complexity. Interns don’t let you fool yourself.

The Buddy System Trap

Let me be clear: don’t abandon buddy systems. They’re valuable. A human being who can answer questions, provide context, and make someone feel welcome is always going to be part of good onboarding.

But buddy systems can also mask the problem. When your new hire’s buddy spends three hours walking them through how to troubleshoot a thousand lines of NX config, that feels like mentorship. It looks like the system working. In reality, it’s a human being absorbing the cost of complexity that shouldn’t exist. How do you measure the extra load being placed on that buddy? How do you account for the work they’re not doing while they’re teaching someone to navigate around problems that could be eliminated?

The buddy isn’t a solution. The buddy is a painkiller. And painkillers are great — right up until you use them as a reason not to fix the underlying injury.

The F5 Experience

Those same Readify consultants taught me something else I’ve turned into a persistent, maybe slightly annoying, initiative. I call it the F5 Experience. Back when we were all using Visual Studio, F5 was the shortcut key for “Start Debugging.” The idea is simple: setting up a repository should require exactly three steps. Clone the repo. Open it in your IDE. Press F5.

Anything more than that, and you have things to fix.

I wrote about this in more detail in a previous post, but the core principle bears repeating: every additional setup step is a documented failure of your developer experience. And the troubleshooting guide is even more revealing. When I sit down with engineers and ask them to document the steps to set up their repo and run it locally, plus a troubleshooting section for common problems, the result tells me everything I need to know.

If you end up with a seven-page Confluence article for setup and another ten for troubleshooting, congratulations — you’ve just documented your problems. Now fix them.

Waste Elimination Starts with “Do We Need This?”

Here’s where lean manufacturing has something to teach us. In the Toyota Production System and its intellectual descendants, the first step in process improvement isn’t “how do we do this faster?” It’s “do we need this at all?”

Waste elimination. Before you optimise, question existence.

If your troubleshooting guide has ten extra steps for the testing DSL you’re using, the first question isn’t “how do we make these steps clearer?” The first question is: do we need that testing DSL? What value is it actually bringing? Is the cognitive overhead it creates for every new person who touches the system justified by the benefits it provides?

Be brutal in your questioning. Because this is what your engineers need — someone willing to ask the hard questions that they didn’t ask, or were too deep in the existing paradigm to think of asking. Every tool, every abstraction, every layer of indirection should earn its place. If it can’t, remove it. Deleted complexity is always better than documented complexity.

As the management theorist Peter Drucker observed, “There is nothing so useless as doing efficiently that which should not be done at all.” That testing DSL you’ve painstakingly documented? That elaborate local environment setup with Docker Compose orchestrating nine services? Question whether they need to exist before you invest in making them easier to use.

Onboarding Isn’t Just for New Hires

Here’s the part that most organisations miss entirely: onboarding doesn’t only happen when someone new joins the company. It happens every time a team needs to contribute to a system they don’t own.

If you’re pursuing feature teams — or stream-aligned teams, depending on your preferred framework — your engineers are constantly onboarding into unfamiliar parts of the codebase. Every cross-team contribution is a mini onboarding event. Every pull request into another team’s service requires enough context to make a meaningful change without breaking something.

This is why onboarding quality matters far beyond HR metrics and time-to-first-commit dashboards. It directly determines whether your organisation can function as a set of cross-functional teams or whether it collapses back into siloed component teams because the cost of contributing across boundaries is simply too high.

The easier you make onboarding to any part of your system, the more fluid your organisation becomes. The harder it is, the more your teams calcify around the code they already know, and the more you end up with single points of failure — those engineers who are “the only person who knows how that service works.”

The Hidden Cost Nobody Tracks

Companies measure time-to-first-commit. They track whether a new hire shipped code in week one. They don’t track time-to-confident-contribution — the point where someone understands not just what they’re shipping but why, where they can make architectural decisions rather than just following instructions, where they stop feeling like an impostor navigating someone else’s system.

That gap — between first commit and confident contribution — is where engineers either click or quietly start updating their CV. And the attrition cost of the ones who leave? It never shows up in any dashboard. It’s invisible, which means it’s easy to ignore, which means it persists.

The engineers who struggle silently for months before either ramping up or giving up represent an enormous hidden cost. Not just the recruitment expense of replacing them, but the opportunity cost of all those months where they could have been productive if the system had been designed to let them be.

The Bottom Line

Onboarding isn’t an HR problem. It’s not a documentation problem. It’s an architecture problem and a developer experience problem wearing an HR costume.

The reason your new engineer can’t be productive in week two isn’t because your wiki is disorganised — though it probably is. It’s because your system is so complex, so dependent on tribal knowledge, and so burdened with accidental complexity that no amount of onboarding documentation can compensate for what is fundamentally a design failure.

The good news? Every improvement you make to onboarding — every setup step you eliminate, every unnecessary tool you remove, every piece of tribal knowledge you codify into the system itself — benefits everyone. Not just new hires. Not just interns. Your senior engineers too, the ones who’ve been quietly absorbing that cognitive load so long they’ve forgotten it’s there.

Can a new joiner push code to production on day one? Maybe not yet. But that’s the nirvana state. That’s the kaizen goal — continuous improvement, always moving toward it, knowing you might never fully arrive but measuring your progress by how close you get.

Now, if you’ll excuse me, I need to go review our intern programme material for this summer. We got onboarding down to a week last year. I’d like to see if we can do it in four days.