A postmortem can be factually correct and still change nothing. Configuration error appears as the cause, better review appears as the action, and the next engineer faces the same interface and the same missing guardrail. The document described the event but did not alter the system.

Reconstruct what was knowable at the time

Avoid hindsight. Build the timeline from logs, chat and interviews, including missing signals and competing alerts that shaped the decision.

Keep contributing conditions

A bad setting may combine with absent validation, broad rollout, weak alerts and slow access. Reducing the story to one root cause removes opportunities to limit the next incident.

Replace reminders with controls

Training can help, but repeatable risk needs automated validation, smaller blast radius, earlier detection or a faster recovery path. Pair a quick mitigation with longer structural work.

Verify the effect

Insert the previously bad configuration and confirm the pipeline blocks it. Give the revised runbook to someone new and run the recovery. An action closes when behavior changes, not when a ticket is marked done.