On-Call Shouldn’t Feel Like Punishment
Or: Why Your Engineers Treat Production Alerts Like Car Alarms in a Parking Lot
It was 2:14 AM on a Wednesday when the PagerDuty notification lit up the bedroom. The on-call engineer glanced at the screen, watched the alert for about ten seconds, then set the phone face-down on the nightstand. Fifteen minutes later, the alert auto-resolved. He rolled over and went back to sleep. When I asked him about it the next morning, he shrugged: “There’s no alert anymore, so it’s not a problem.”
That sentence should terrify every engineering leader reading this. Not because the engineer was lazy — he wasn’t. He’d been on rotation for months. He’d been woken up dozens of times for alerts that resolved themselves before he could even open his laptop. He’d investigated spikes that turned out to be nothing, chased phantom errors through dashboards that told him what was happening to servers but never what was happening to customers. He’d learned — through repeated, exhausting experience — that most of these alerts meant nothing. No context on business impact. No way to tell if a single user was affected or ten thousand. Just a graph going red, then going green, then going red again next Tuesday.
So he did what any rational person would do. He waited it out.
That’s the thing nobody wants to admit: he wasn’t ignoring the alert. He was responding perfectly logically to a system that had taught him, over and over, that alerts and customer impact were two completely unrelated things. We’d built an on-call culture where production incidents existed in a vacuum — disconnected from the people actually using our product — and then acted surprised when engineers stopped treating them as urgent.
As the quality pioneer W. Edwards Deming put it, “A bad system will beat a good person every time.” We hadn’t hired the wrong engineers. We’d built the wrong system — one where neither the team nor the individual had the tools or context to respond meaningfully. Just the dull weight of a rotation nobody wanted and alerts nobody trusted.
The Canary in the Coal Mine
You already know if your on-call culture is broken. The warning signs aren’t subtle — they’re screaming at you from your rotation schedule.
The first thing that happens is the rotation length starts shrinking. What was once a week becomes three days. Then two. Then daily. And daily rotations are the organisational equivalent of a hot potato — you’re not managing incidents, you’re managing the handoff. Every twenty-four hours, context evaporates. That flaky service that started misbehaving at 9 PM? By the time the next engineer picks up at 9 AM, it’s someone else’s archaeology project. Nobody investigates because nobody owns it long enough to care.
The second signal is the wait. When alerts fire and engineers just… watch. They’ve learned, through bitter experience, that most alerts resolve themselves within fifteen minutes. So they sit. They watch the notification. They count. If it goes away, they go back to whatever they were doing. If it doesn’t, well, maybe it’s a real problem. Maybe.
The psychologist Martin Seligman called this “learned helplessness” — the state where repeated exposure to uncontrollable situations causes an organism to stop trying entirely. His experiments involved dogs and electric shocks, but the mechanism is identical in on-call rotations. When engineers can’t distinguish between noise and signal, when they can’t connect alerts to real impact, they stop responding. Not out of malice. Out of rational self-preservation.
And here’s the part that really stings: when you ask them about it — when you sit down and say “are you going to investigate that?” — the answer reveals everything. “There’s no alert anymore, so it’s not a problem.” They’re not wrong within the framework we’ve given them. If the only definition of “problem” is “active alert,” then a resolved alert is, by definition, not a problem. We taught them that.
The Ownership Gap
Let’s dig into the root cause, because it’s not laziness and it’s not incompetence. It’s disconnection.
When engineers can’t connect the work they do day-to-day with the impact on actual users, they stop caring about the impact on actual users. That sounds harsh, but it’s not a moral judgment — it’s a predictable consequence of how we’ve structured their work. They’ve stopped thinking about customers and started thinking exclusively about the next feature they’re shipping. Production incidents become interruptions to their “real” work rather than the most direct feedback loop they have to the people actually using what they build.
This is a failure of management. Full stop. I’ve had to recover many a team from this exact situation, and it always traces back to the same root: we didn’t build a culture where customer impact was real, visible, and personally felt. We talk about “customer first” in product engineering — it’s practically tattooed on the office walls — but if your engineers can’t tell you how a 500-error spike at 2 AM affects the person trying to book a hotel room, those words are just decoration.
Here’s a question I regularly pose to my engineers: “If a server is at 100% CPU at 2 AM, should you get woken up to fix it?”
For most of them, this is actually a pretty easy question to answer: “Not if there’s no business impact.”
Good. They understand the principle. So why are we waking them up for it anyway?
Alert Fatigue Is a Design Problem, Not a People Problem
This is what I often find at the heart of on-call apathy: engineers drowning in alerts they can’t contextualise. They see a traffic spike. A CPU graph going red. A latency percentile breaching a threshold. But what real impact does this have? Without being able to answer that question, every alert sounds exactly the same — which is to say, they all sound hollow.
Alert fatigue isn’t solved by tuning thresholds, though that helps. It’s solved by connecting alerts to outcomes. When an engineer gets paged at 2 AM, they need to be able to answer one question immediately: “Is this hurting our customers right now, and how badly?”
The sports analyst and statistician Bill James — the man who revolutionised baseball by insisting that what gets measured should actually matter — once said, “People want to be given credit for the things they’ve done. You have to measure what actually counts.” We’ve been measuring CPU utilisation and error rates when we should be measuring customer experience and business impact. The metrics aren’t wrong, but they’re incomplete in a way that renders them meaningless to the person staring at them at 2 AM.
This is where north star metrics become your most powerful tool in fixing on-call culture. North star metrics are the common currency an organisation uses to define success — the numbers that product, engineering, and leadership all agree actually matter. Your features from product should be measured against these metrics. Your engineers shouldn’t be building features; they should be building things that move these numbers. And this is how you measure them.
I wrote about this approach in more detail in my earlier post on experimentation and building a culture that bets on better. The core idea is simple: if everyone in your organisation is oriented around the same definition of success, the conversations change. “Should we fix this flaky service?” stops being a negotiation and starts being an obvious yes, because everyone can see it degrading the metric they all care about.
The Common Currency That Changes Everything
Follow through on this. If you can get to the point where you’re able to measure a drop in your north star metrics using semantic monitoring — I wrote about this approach here — then you can attribute value to production health easily. And this has a huge impact on engineers.
When a 2 AM alert comes with context that says “conversion rate has dropped 3% in the last fifteen minutes affecting an estimated X bookings,” you’ve transformed that alert from noise into narrative. The engineer isn’t responding to a graph anymore. They’re responding to real people having a degraded experience. That’s not a subtle distinction — it’s the difference between an alert that gets ignored and one that gets investigated.
You will create a common currency spoken by both product and tech. Product cares about conversions and revenue. Engineering cares about system health and code quality. North star metrics, measured through semantic monitoring, bridge that gap. When both sides are looking at the same numbers, the apathy dissolves — because the work stops being abstract and starts being consequential.
And yes, there is a direct, proven link between page performance and conversion. We’ve demonstrated it. A host of other major companies have demonstrated it. When your site is slow, when it’s unreliable, when it’s “wigging out” on customers while they’re trying to complete a purchase — they leave. They don’t file a bug report. They don’t send you feedback. They just leave, and they take their money with them. Web vitals aren’t vanity metrics; they’re revenue metrics. When engineers understand that, the on-call conversation changes fundamentally.
Rebuilding the Culture: What Actually Works
So you’ve diagnosed the problem. Your on-call is a punishment rotation, your engineers have learned helplessness, and your alerts are disconnected from anything that matters. Now what?
Make On-Call Educational, Not Extractive
The worst on-call rotations are purely extractive — they take energy and time from engineers and give nothing back. The best ones are educational. Every incident is a window into how your system actually behaves under stress, which is information you can’t get from reading architecture diagrams in a comfortable meeting room.
Post-incident reviews should be blameless, genuinely curious, and focused on system improvements rather than individual failures. When an engineer spends a night debugging a cascading failure, that experience should be captured and shared, not just endured and forgotten. The goal is to make each incident make the system — and the team’s understanding of it — better.
Connect Alerts to Impact, Not Symptoms
Rearchitect your alerting to answer “what’s happening to customers?” before “what’s happening to servers?” A CPU spike isn’t inherently a problem. A 2% drop in successful checkouts is. When your alerting speaks in customer impact, engineers can make rational decisions about urgency — and rational responses to meaningful alerts look very different from learned helplessness in the face of noise.
Reduce the Burden Over Time
Here’s the litmus test for a healthy on-call culture: the number of pages should be decreasing over time. If every quarter your team gets paged just as often as the last, you’re not running an on-call rotation — you’re running a suffering rotation. Each incident should produce a concrete action that prevents recurrence. Track pages-per-rotation as a team metric and make reducing it an explicit goal.
The best teams I’ve worked with treat every page as a small failure of the system — not of the person who gets paged, but of the engineering decisions that allowed the page to happen. That framing turns on-call from a burden into a feedback mechanism.
Own the Rotation Length
Stop letting rotation length shrink as a coping mechanism. If your engineers are reducing rotations to daily handoffs, that’s not a scheduling preference — it’s a symptom of a broken system. Fix the system. Make on-call manageable enough that a week-long rotation doesn’t feel like a prison sentence. That means fewer alerts, better runbooks, clearer escalation paths, and — critically — the authority to actually fix things rather than just acknowledge them.
The Leadership Obligation
Let me be direct about something: if your on-call culture is toxic, that’s on you. Not on the engineers, not on the SRE team, not on the alerting tools. On you, the engineering leader.
Without customers, we are nothing. That’s not a motivational poster — it’s a statement of economic reality. Customers care about their experience. They care about quality, performance, and a product that works when they need it. Part of our job as engineering leaders is building a culture where that reality is felt, not just stated. Where production health is everyone’s concern, not just the unfortunate person whose name happens to be on the rotation this week.
As Simon Sinek put it, “The goal is not to be perfect by the end. The goal is to be better today.” You’re not going to fix on-call culture in a quarter. You’re not going to eliminate all false alerts by next sprint. But you can start connecting alerts to business impact today. You can start running blameless post-mortems this week. You can start measuring pages-per-rotation this month.
The Bottom Line
On-call shouldn’t be a punishment. It should be a first-class engineering concern — a design constraint baked into how we build systems, not an afterthought bolted on after deployment. The teams that get this right don’t just have happier engineers. They have more reliable systems, faster incident response, and — here’s the part that gets leadership’s attention — better business outcomes.
The path from “on-call as suffering” to “on-call as ownership” runs directly through measurement. North star metrics give everyone the same definition of success. Semantic monitoring makes production health visible and attributable. Customer impact framing turns abstract alerts into meaningful signals. And blameless incident culture turns every page into an investment in a better system.
Your engineers aren’t apathetic by nature. They’re apathetic because we built systems that taught them apathy was rational. Fix the systems, connect the dots between alerts and impact, give them a common language shared with product — and watch what happens when on-call becomes something engineers actually learn from instead of merely survive.
Now, if you’ll excuse me, I need to go investigate why my own team’s alerting dashboard has three amber warnings that have been “about to resolve themselves” since yesterday morning. Physician, heal thyself.