Zavodit Авг 6, 2026 7 мин

Blameless Post-Mortem Engineering: How Startups Should Run Incident Reviews

How to run a blameless post-mortem that produces action items engineers actually implement. A practitioner's template and cultural guidance from 200+ startup projects.

A
Aleksandr Protsiuk Fractional CTO - Саннивейл, Калифорния
Опубликовано 06.08.2026 Обновлено 06.08.2026 Время чтения 7 мин
CTO

I joined an engagement two weeks after a startup had its first significant outage. Five hours of downtime, a SaaS product, during business hours. Customers had noticed. Two had sent formal complaints. One had asked about SLA credits.

The team had done a post-mortem. I read the document. It had a timeline of events, a root cause identified as "database connection pool exhausted," and a single action item: "investigate connection pool settings."

Six weeks later, the action item was still open. There had been two more incidents, both related to the same underlying problem.

The post-mortem had described what happened without understanding why it happened. It had produced an action item that was too vague to assign, too vague to complete, and too vague to close. The team had gone through the motions of a post-mortem without getting any of the value.

This is how most post-mortems fail.

What a Post-Mortem Is Actually For

The purpose of a post-mortem is to prevent the next incident, not to document the last one. This sounds obvious but it is not - the way most post-mortems are run produces documentation without prevention.

Prevention requires two things that documentation does not require:

Understanding the system failure deeply enough to know what would have to change to prevent recurrence. In the database connection pool case: why did the pool exhaust? Because traffic spiked? Because a query was slow? Because a connection was leaking? Each answer points to a different fix. "Connection pool exhausted" is the symptom, not the cause.

Action items that are specific enough to be completed and assigned to owners. "Investigate connection pool settings" is not an action item. It is a prompt for further investigation that could lead to an action item. The post-mortem is the place to do the investigation, not defer it.

The blameless principle - running a post-mortem without assigning blame to individuals - is not softness. It is a practical necessity. If post-mortems result in people being blamed, people stop being honest in post-mortems. You get sanitized timelines and root causes that attribute the problem to an abstract system failure rather than the human decisions that contributed to it. Blameless means you can surface the real sequence of events and decisions without anyone defending themselves.

The Template That Works

This is the structure I use for every post-mortem, regardless of incident type.

Incident summary. Two to three sentences. What broke, when, for how long, what the user impact was. This section should be completable before the post-mortem meeting so everyone comes in with shared context.

Timeline. A chronological sequence of events, starting before the first symptom. Include: when the problem started, when it was detected, what triggered detection (alert, user report, engineer noticed), what actions were taken and by whom, when the problem was resolved. Be specific about times. "Around 2 PM" is not a timeline entry. "14:07 - first alert fired" is.

Root cause. This is the hardest part and the most often done poorly. The method I use is five whys: ask "why did this happen?" five times, each answer becoming the premise of the next question.

Example:

The last answer is the actual root cause. The fix is adding database query performance review to the PR checklist - not adding an index (that is the fix for this specific incident, but it does not prevent the next missing index).

Contributing factors. Things that made the incident worse without being the root cause. In this example: monitoring did not alert until five minutes after the problem started, which extended the time to detection. The on-call rotation meant the person who knew the database best was not the one who first responded.

What went well. What worked during the incident response? This is not optional padding. It identifies practices to preserve and builds the morale of teams who responded well under pressure.

Action items. Each action item must have: a description specific enough that someone can know when it is done, an owner (a person, not a team), and a deadline. "Add database query performance review to PR checklist" owned by [specific engineer] due [specific date] is an action item. "Improve monitoring" is not.

The Meeting That Produces the Document

The post-mortem document is produced in a meeting that typically runs 60-90 minutes. How you run this meeting determines the quality of the output.

Set the tone at the start. State explicitly: we are here to understand the system failure, not to assign blame. If an individual made a decision that contributed to the incident, we want to understand what information they had at the time and what system or process factors led to that decision. No one is on trial.

Build the timeline collaboratively. Have people add to the timeline in real time - the person who noticed the first alert, the person who made the first change, the person who found the actual problem. Different people have different parts of the story. The collaborative timeline is always more complete than any individual's account.

Do the five whys with the group. Do not let one person drive the root cause analysis. Ask the questions aloud and let the group respond. People will disagree about why something happened, and that disagreement is valuable - it surfaces different mental models of how the system works.

Write action items during the meeting, not after. If an action item is identified, write it in the document before moving on. Assign it to someone in the room, negotiate the deadline in the room. Action items that get written "after the meeting" often do not get written at all.

The Follow-Up That Most Teams Skip

A post-mortem is complete when the action items are done, not when the document is written. Most teams do the document and not the follow-up.

The follow-up requires a process: action items from post-mortems go into the team's tracking system (Jira, Linear, GitHub Issues) the same day the post-mortem runs. They are reviewed in the weekly engineering meeting. Incomplete action items from a post-mortem are a standing agenda item until they are done.

The signal that a post-mortem culture is healthy: action items get done before the next incident, not after.

The signal that it is not: the same class of problem causes an incident twice in three months. This almost always means the previous post-mortem produced action items that did not get implemented.

Building a Post-Mortem Culture From Scratch

If your team has never run a post-mortem, or has run them without results, the first thing to establish is the norm: post-mortems are expected for any significant incident, and they are completed within 48 hours of the incident being resolved.

The 48-hour window matters. After 48 hours, memories are less precise, the team has moved on, and the details of what happened become harder to reconstruct. Post-mortems run a week after the incident are post-mortems in name only.

Start with smaller incidents if major ones are rare. A post-mortem practice that only runs after catastrophic failures is a practice that runs twice a year. Teams learn incident response by running the process repeatedly. Run post-mortems on significant performance degradations, near-misses, and customer-reported bugs, not just full outages.

The goal is a team that treats every incident as a learning opportunity and has a reliable process for extracting and implementing the lessons. That culture does not appear spontaneously. It is built deliberately, with consistent leadership over six to twelve months.

Book a 30-minute call: https://calendly.com/alpsf/zoom-with-aleksandr

Теги

Было полезно? Поделитесь.

A
Aleksandr Protsiuk
Fractional CTO - Саннивейл, Калифорния

15+ лет в разработке. 200+ продуктов. Победитель APIWORLD 2024 Hackathon в Silicon Valley. Работаю как fractional CTO для стартапов -- архитектура, AI-first разработка, найм, техническое due diligence.

Рассылка - подписка

Каждый выпуск -- к вам на почту.

Одна большая статья в неделю. Без спама, без SEO-воды. Пишет практикующий CTO, который все еще шипит код.

Подписаться - Отписка в один клик
Подписка оформлена

Добро пожаловать. Скоро напишем.