Beer & Servers Don't Mix

The Autopsy Nobody Wants to Do

Or: A Field Guide to the Execution Post-Mortem

Somchai was the first one in the room. It was 10:14 on a Thursday morning and the air conditioning on level 6 was doing its usual impersonation of effort — technically running, technically failing. He’d brought his laptop. He didn’t open it. He just sat at the end of the table and looked at the door, the way you look at a door when you know who’s coming through it and you know what they’re going to say. The room smelled of the Thai tea he hadn’t touched. The calendar invite had said “Milestone Review.” Everyone in the meeting knew that was a polite fiction. What it meant was: the milestone was missed, and now we’re going to talk about it.

That meeting has a specific atmosphere. You’ve been in it. You know the particular quality of silence before it starts — not the comfortable silence of people thinking, but the loaded silence of people preparing. The product side rehearsing their version. The engineering side marshalling their context. Everyone in the room carrying a narrative they’ve been quietly sharpening since the milestone slipped.

What happens next is almost always the same. Someone says “failed.” Someone else says “it was complicated.” Someone mentions a dependency. Someone mentions the requirements changing. Someone’s voice tightens a little. Nobody’s lying, exactly — but nobody’s right, either. Because what’s happening isn’t a conversation about what happened. It’s a competition between memories.

This post is about how to stop that competition before it starts. Not by being a better mediator. Not by having better values or running better retros. But by walking into the room with something that memories can’t compete with: evidence.

The Problem With the Conversation We Keep Having

Most delivery post-mortems — when they happen at all — are conducted from memory. People describe their experience of the sprint. They talk about what felt slow, what felt blocked, what they were waiting on. And because everyone experienced the same period of time from a different vantage point, you get a beautiful collection of individually coherent, mutually incompatible accounts.

The loudest voice tends to win. The most senior person’s narrative tends to stick. The engineer who was quietly blocked for two weeks — blocked in a way that never made it to a standup, never landed in a ticket comment, never got escalated — stays quiet, because what are they going to say? “I was waiting”? Against someone with a slide deck?

Here’s the part that should genuinely unsettle you: in my experience running these investigations, the thing being blamed loudest in the room is almost never the actual constraint. The finger points at QA. The timeline shows QA was waiting two days for a build. The finger points at another team’s API. The timeline shows the dependency wasn’t tracked as a risk at planning and the team kept working as if it didn’t exist. The finger points at scope creep. The timeline shows the scope was defined precisely — the estimate was just wrong, and nobody re-evaluated it when reality diverged.

The culprits are almost always quieter. Internal silos. Analysis paralysis that looked like thoroughness. Large PRs opened after weeks of solo development that assumed alignment that hadn’t been established. Poor planning that the team knew was poor by day three but kept pushing through anyway, hoping.

You can’t see any of that in a retro. A retro is the wrong tool for this problem. A retro is two weeks of memory, fifteen post-it notes, and a timebox. A milestone failure usually spans multiple sprints and requires you to go back further than anyone’s accurate recall allows. What you need isn’t a retrospective. It’s a reconstruction.

That’s what I call the Execution Post-Mortem.

What the Execution Post-Mortem Is

An Execution Post-Mortem is a structured, evidence-based investigation into how a milestone actually unfolded — built from primary sources rather than memory.

Not feelings. Not narratives. Primary sources: ticket history, git logs, PR review timelines, deployment records, calendar data. Things that were written down at the time. Things that don’t shift depending on who’s telling the story.

The output is a Delivery Timeline: a chronological record of what happened to each piece of work, from the moment it entered the system to the moment it was done (or wasn’t). It shows you where time actually went. Day by day. Gap by gap.

Once you have the timeline, the conversation in that meeting room changes completely. Instead of “it was complicated,” you can say: “Three stories sat in review for more than five days each — here’s what was happening during those gaps.” Instead of “the requirements kept changing,” you can say: “The story acceptance criteria changed after pointing, on this date, and we didn’t re-size or flag it — here’s what we’re doing differently.”

The timeline doesn’t assign blame. It assigns ownership. And ownership is something both product and engineering can actually work with.

As the journalist and author Katherine Boo once wrote about investigating complex human systems: “The building of a narrative requires the destruction of comfortable assumptions.” That’s exactly what a delivery timeline does. It destroys the comfortable assumption that you know what happened — and replaces it with what actually happened.

Why It Works Psychologically

Before we get to the steps, this is worth understanding: the method works not just analytically, but socially.

When you walk into that room with a timeline rather than a narrative, you change the dynamic. You’ve made the problem visible to everyone simultaneously, from the same data, at the same time. Nobody finds out what happened by listening to someone else’s account — everyone looks at the same picture together. That shift — from testimony to evidence — is what lets you have a collaborative conversation instead of a defensive one.

The young backend engineer who moved on to the next milestone while the frontend work sat with one person for three weeks? When the timeline surfaces that, his reaction isn’t defensiveness. It’s surprise. Because he genuinely didn’t know. He thought moving ahead was being efficient. Nobody had shown him otherwise. He was output-focused, not outcome-focused — and nobody had ever told him those two things could diverge. The timeline made the invisible visible, and the conversation that followed was coaching, not confrontation.

That’s the version of this meeting worth engineering toward.

But here’s the caveat, and it matters: the method is only as good as the conditions it’s run in. If the room isn’t psychologically safe, the same data that produces honest conversation in one place produces counter-accusations in another. The timeline is neutral. The people in the room are not. Your job, as the person running this, is to establish the conditions before the first data point goes on the board.

Running It Blameless: What That Actually Means

“Blameless” gets thrown around in engineering culture as if it’s self-explanatory. It isn’t. Here’s what it means in practice when you’re running an Execution Post-Mortem.

No names on the board. When you’re building the timeline and presenting findings, you talk about the story, the PR, the sprint — never the engineer. Not “Somchai’s PR sat for five days” but “this PR sat in review for five days.” Not “the backend team didn’t flag the dependency” but “this dependency wasn’t flagged at planning.” The moment a name goes up, the person attached to it stops thinking about the problem and starts thinking about themselves. Everyone else in the room does the same.

“We” owns the process. Nobody owns the mistake. The framing is always: what did our process allow to happen? A story sat unreviewed for five days because we don’t have a review SLA. A dependency was invisible because we don’t track them explicitly at planning. An engineer carried work alone for three weeks because we don’t have a practice of checking in on single-assignee stories. The team built the conditions. The team changes the conditions.

Findings are systemic, not personal. The question is never “who made this decision?” It’s “what made this decision rational at the time?” Because almost every decision that looks wrong in hindsight looked reasonable when it was made. The engineer who built solo for three weeks wasn’t being reckless — he thought he was being efficient. The context that made that feel right is the thing worth examining.

None of this means individuals are never accountable. It means accountability lands on the right thing: what does this person need to do differently, and what does the system need to give them to make that possible? That’s a coaching conversation. It happens separately, privately, and after the post-mortem — not in the room.

When the Room Goes Off-Script

Even with all of that set up in advance, rooms don’t always cooperate. People are human, and delivery failures carry real frustration. Sometimes someone fixates on a person instead of a process and the energy in the room shifts from analysis to prosecution.

I’ve been in one of those. A very animated Italian engineer — loud, passionate, entirely convinced that the problem was a specific colleague who had made a mistake. He wasn’t wrong that a mistake had been made. He was completely wrong about why it mattered. He kept circling back to the individual, louder each time, while the actual process failure sat quietly in the data, unexamined.

I let him run for a bit. Then I said, very calmly: “Don’t worry about him. After this meeting, we’re going to take him out the back and shoot him.”

He stopped mid-sentence. Then he started laughing — genuinely, loudly, the kind of laugh that releases something. When he came back up for air, he’d lost the thread of his rant. He looked slightly sheepish. And then we got back to the process.

I’m not recommending gallows humour as a universal facilitation technique. But the underlying move is worth understanding: you need a pattern interrupt, something that breaks the emotional momentum and creates a beat of distance. Humour works when the room already trusts you. A direct redirect works in other contexts — “I hear you, and that’s a separate conversation. Right now I want to understand the system. Can we come back to the timeline?” The specific tool is less important than the intent: you’re not dismissing the frustration, you’re redirecting the energy toward something that will actually produce a result.

The alternative — letting the room stay in blame mode — is how you end the meeting with one person feeling vindicated and everyone else feeling worse, and the process failure completely unaddressed.

The Process: Step by Step

Step 1: Establish Your Anchors

Start with two dates you can verify from a primary source — not from memory, not from “I think it started around”:

  • T₀ — The date the first story for this milestone was added to a sprint. Not when the initiative was announced. Not when it was discussed in planning. When did it first land in an active sprint backlog?- T_end — Either the actual completion date, or the date it became objectively clear it wouldn’t complete on time. Write these down. Every other data point you collect goes between these two anchors. You now have the edges of your picture.

Step 2: Map the PR Lifecycle for Every Story

This is your primary data source. The git history, PR review history, and deployment record are all visible from the PR itself — and a PR has four dates that will give you 95% of the story:

  • First commit — when development actually started- Ready for review — when the engineer considered it done and opened it up- Approved — when the team signed off- Merged / in production — when it actually shipped For a multi-sprint milestone you may be looking at a lot of PRs. Don’t try to investigate all of them at this stage. Just collect these four dates for each one, put them in a row, and calculate the gaps. That’s enough to see the shape of the problem. The gaps that stand out are the ones worth drilling into.

What you’re looking for at a glance:

  • A long gap between first commit and ready for review — was the engineer working in isolation for too long before seeking feedback?- A long gap between ready for review and approved — where was the bottleneck: waiting for a reviewer, review rounds, rework?- A long gap between approved and merged/production — change windows, deployment failures, manual steps that shouldn’t exist? Most PRs will look fine. A few will have gaps that jump off the page. Those are the ones that need a second look. The four-date summary tells you which ones. Only then do you go deeper — into review comments, commit frequency, deployment logs — on the ones that earned it.

A note on spike stories: not everything in your milestone will have a PR. Spikes — investigation, design, and research stories — produce a document or a decision, not a commit. Don’t skip them. In fact, treat them with extra attention, because spikes are where analysis paralysis hides most comfortably. A spike that was estimated at two days and ran for two weeks won’t show up in any PR scan. It’ll just be a ticket that sat “In Progress” for a very long time while the team convinced themselves they were being thorough. For spikes, use the ticket history directly: when was it picked up, when was it closed, and what came out of it? If the answer to that last question is “another spike,” you’ve found something worth discussing.

Step 3: Pull the Ticket History for Context

The ticket history fills in what the PR can’t tell you: what happened before the first commit, and anything that happened outside the code.

For each story:

  • When was it created in the backlog?- When was it added to the sprint?- When did it move to “In Progress”?- Any status reversals — stories moved back to “To Do” mid-sprint?- Any comments flagging blockers, dependency issues, or scope changes?- Any assignee changes? The gap between “added to sprint” and “first commit” is the one to watch here. It’s often where planning failures live — a story that sat in a sprint for four days before anyone touched it, because it wasn’t actually ready to be worked on, or because the person who picked it up had three other things running.

Most Jira instances keep a full audit trail. Pull it. The history was written at the time and doesn’t shift depending on who’s in the room.

Step 4: Build the Table

Take every data point and lay it out. For each story, you want the ticket anchor dates alongside the four PR lifecycle dates, with the gaps calculated between each. It should look something like this:

If you have a whiteboard, draw it up as an actual timeline though, visualizing it like this for everyone to collaborate on works wonders, also using distance a time, its easy to compare, but make sure your scale is correctly proportionated to the time dimension, otherwise you’ll mask problems potentially.

UNF-114 and UNF-122 look fine. UNF-118 jumps off the page immediately. But so does UNF-116 — a spike that ran for sixteen days with no PR, no commit, nothing. The ticket just sat open. Whatever it was investigating, it took four times longer than it should have, and the stories that were presumably waiting on its outcome couldn’t start until it closed. That’s analysis paralysis made visible.

That’s your investigation target. Now you go deeper on those two — into the PR comments, the ticket history, the calendar — and leave the others alone.

Total calendar days for UNF-118: 27. Original estimate: one sprint. Total calendar days for UNF-116: 16. Original estimate: three days.

The question isn’t “why did these take so long?” in the abstract. For UNF-118 it’s three specific questions: why four days before the first commit? What happened in review for seven days? And what was blocking the merge for eleven days after approval? For UNF-116 it’s simpler and harder: what were we actually doing for sixteen days, and what decision came out of it?

That’s the investigation. The table doesn’t give you answers. It tells you which questions to ask — specific, answerable ones, grounded in data rather than retrospective impression.

Step 5: Interrogate the Gaps

For every gap that looks disproportionate relative to the story’s estimate, ask: what was happening on those days? A two-day gap on a one-point story is worth a question. The same two days on a thirteen-point story is probably noise. The threshold isn’t a fixed number — it’s your judgment about what looks out of proportion given what the work was supposed to be.

The timeline flags the gaps. The team explains them. Once you have the table in front of everyone, simply ask: “UNF-116 was open for sixteen days against a three-day estimate — can someone walk us through what was happening there?” You’ll get the answer in thirty seconds, and it will be more accurate than anything you’d reconstruct from Slack history or calendar data. The people who did the work know what happened. The timeline just gives them something specific to respond to, instead of a vague invitation to defend themselves.

Most gaps have an explanation. The question is whether that explanation was visible at the time, or only visible now in retrospect. That distinction is what tells you whether you have a process problem to fix or just a bad run of luck.

Categorise what you find. In practice, across many of these investigations, the patterns that surface most often are:

Poor planning that nobody re-evaluated. The estimate was wrong at the start. The team knew mid-sprint it was wrong. Nobody raised it. The problem compounded silently until the milestone slipped.

No upfront cross-team discussion, then a giant PR. An engineer goes heads-down, builds something substantial, opens a large merge request — and gets significant rework feedback. The time lost isn’t in the code. It’s in the weeks of parallel work that assumed alignment that was never established.

Analysis paralysis. The overcorrection to the above. Extensive design documents, lengthy alignment meetings, decision-by-committee before any code is written. A different kind of delay, but the timeline catches it just as clearly. The pendulum swings both ways.

Internal team silos. The most insidious, because the team often doesn’t realise it’s happening. Three weeks of frontend work sitting with one engineer while the backend engineers consider themselves done and move to the next milestone. Nobody raised a flag. Nobody re-assigned. The timeline shows a single assignee on a story for three weeks with no collaboration signals. Nobody looks negligent. They just weren’t paying attention to the right thing.

And here’s what makes it a leadership failure as much as a process one: the backend engineer who moved on wasn’t being careless. He thought he was being efficient. Nobody had ever told him it was acceptable — encouraged, even — to slow down and help with the frontend work, to go a bit slower individually so the team could reach the goal together. Nobody had told him to optimise for outcome, not output. He was measuring himself by what he shipped. The team was supposed to be measured by what they delivered. Those two things had quietly diverged, and the timeline is what made it visible.

And the thing that gets blamed loudest but is rarely the actual constraint: QA. QA teams tend to be efficient partly because they know they attract blame and compensate accordingly. If the timeline surfaces QA as a genuine bottleneck, it’s worth noting — but be prepared for it to point elsewhere first.

External team dependencies. This one is common enough to name explicitly, because it will surface in almost every multi-team milestone — a story that couldn’t move because it was waiting on another team’s API, another team’s capacity. It shows up in the timeline as a clean gap: work picked up, nothing happened, work resumed. The team will be able to tell you exactly what they were waiting for.

Here’s the important thing: if this keeps appearing, it is not a problem the team can fix. It’s an organisational design problem. Teams that are structurally dependent on other teams to deliver will keep showing this pattern no matter how well they run their own process. The only durable solution is to organise around value streams — teams that own the full slice of capability they need to deliver, without hard dependencies on others for their day-to-day work. If you keep seeing external dependency gaps in your timelines, that’s the signal worth escalating.

The concepts of stream-aligned teams and feature teams — teams that own the full slice of capability they need to deliver, without hard dependencies on others for their day-to-day work — are two names for essentially the same idea. Team Topologies by Matthew Skelton and Manuel Pais is the clearest framework for thinking about this from an organisational design perspective. LeSS (Large-Scale Scrum) by Craig Larman and Bas Vodde arrives at the same place from a Scrum scaling angle. Accelerate by Nicole Forsgren, Jez Humble, and Gene Kim provides the research evidence for why loosely-coupled team structures correlate directly with delivery performance. If your timelines keep surfacing the same cross-team gaps, those are the places to go next. I also explored this dynamic from a different angle in The Feature Team Fallacy.

Step 6: Write the Findings Document

One page. Two sections.

What happened — the factual timeline summary, the key gaps identified, and what was happening in each one.

What we’re changing — concrete, specific actions. Not “better communication.” Not “clearer requirements.” Things like: “We will add a blocker comment in the ticket within 24 hours of being blocked, rather than surfacing it at standup.” Or: “Stories with external team dependencies will have that dependency explicitly tagged at planning, not discovered mid-sprint.” Or: “Our PR review SLA is 24 hours for first review. We will add a rotation to enforce it.”

Specificity is the test. If you can’t describe what “done” looks like for the action, it isn’t a change — it’s a wish.

Who Runs This, and When

Who: Ideally, the engineering manager. Not because a senior engineer couldn’t do it — they often can — but because the EM is closer to a neutral third party than anyone inside the team. People have their own stake in the outcome, their own version of the sprint, their own relationships with the colleagues whose work is under the microscope. Distance matters when the findings are uncomfortable.

In practice, this process tends to live as tribal knowledge — something a particular EM knows how to do and runs when things go wrong, without it ever being written down or taught. That’s partly why this post exists.

When: Not as a routine cadence. This is a diagnostic tool, not a ceremony. Run it when a milestone is missed or significantly late. Run it when a stakeholder uses the word “failed.” Run it when the same conversation keeps going in circles, when people are defending rather than analysing, when the post-retro action items look identical to last quarter’s post-retro action items. That’s the smell.

The Meeting, Revisited

Go back to that room. Somchai, the untouched Thai tea, the loaded silence.

That silence is energy. It has to go somewhere — into defensiveness, or into analysis. The difference between those two outcomes isn’t people’s intentions or their willingness to be honest. It’s whether there’s something concrete in the room to direct the energy toward.

The timeline is that something. Not because it exonerates anyone, and not because it assigns blame — but because it turns a competition between narratives into a shared investigation. Everyone looking at the same picture, from the same data, at the same time.

When you bring light to the problem, you get collaboration. Nobody wants to have this meeting again next quarter. Nobody is deliberately holding the team back. When the problem is visible and the framing is about process rather than people, everyone works to fix it.

The key is to be the person who walks in with the picture.

The Bottom Line

As Daniel Kahneman observed, “We can be blind to the obvious, and we are also blind to our own blindness.” Delivery failures are rarely caused by bad engineers or bad intentions. They’re caused by things that were invisible — gaps that nobody tracked, blockers that never surfaced, silos that felt like efficiency. The Execution Post-Mortem is a process for making those things visible, systematically and without theatre.

Your retrospective can’t do this. A retro covers two weeks of memory; a milestone failure spans months of reality. Post-it notes don’t surface a three-week frontend silo. Dot-voting doesn’t find your PR review latency.

What you need is the timeline. It’s unglamorous work — ticket exports, git logs, calendar cross-referencing. It takes a few hours. It requires honesty from people who’d rather move on. But it’s the only way to have the conversation that actually changes something.

Now if you’ll excuse me, I need to go pull a ticket history. We missed something last sprint, and I’ve already heard two compelling explanations for why. Neither of them is probably right.