Operating in Production: On-Call, Incident Command, and Reporting Clocks Module 5 · Learn

Read Other People's Postmortems

Last reviewed · content updated

Intermediate

What you'll learn

~20 min
  • Read five providers' own post-incident words for cause and change, without importing a search-summarized paraphrase
  • State what a reader's own incident record would and would not have captured for each case
  • Contrast a utility's multi-month final-report timescale with a same-week corporate write-up

Before the detail — Decision: read five providers’ and one grid association’s own published words about six selected public incidents, and compare what your own incident record would have captured. Outcome: six worked examples of a stated cause—or “cause not stated”—what changed afterward, and the clock each write-up ran on. Artifact: none new — this lesson is a reading protocol applied to the record 1.4 already built. Status of what follows: reusable guidance.

Prompt first: read one, in their words

Here is a public post-incident review from [provider], published
[date] [paste the relevant paragraphs verbatim - no summary].
Extract three things, quoting the source for each:
CAUSE what they say caused it, in their own words
CHANGE what they say they changed afterward, in their own
words
GAP one thing our incident record template (sections 1-10)
would have forced us to write down that their public
write-up does not mention - or "none found"
Do not paraphrase the cause or the change - quote them. If the
write-up does not state a cause plainly, write "cause not stated"
and say what it names instead: a symptom, a mechanism, a hypothesis.

A provider’s postmortem is evidence about a system you do not operate, and the value is in their exact words: this domain has already produced summaries that invented mechanisms and dates no primary source supports, so quoting is the only safe default.

Reading protocol: cause, change, gap

Every case below gets the same three questions. What do they say caused it, in their own words? What do they say they changed afterward, in their own words? And what is one thing your own incident record — the state transitions of section 1, the determinations of section 5, the corrective-action register of section 10 — would have forced someone to write down that their public write-up leaves unsaid? The third question is where the value is: a provider’s postmortem tells you what a competent team chose to publish, not everything their own internal record holds.

Five corporate reviews across three years

Cloudflare, 2025-11-18. Their own words: “a change to one of our database systems’ permissions… caused the database to output multiple entries into a ‘feature file’ used by our Bot Management system. That feature file, in turn, doubled in size.” They are explicit that it was not an attack. What changed, in their words: “Hardening ingestion of Cloudflare-generated configuration files in the same way we would for user-generated input” and “Enabling more global kill switches for features”. The gap a record would catch: their own configuration output was trusted the way user input is not — a determination entry asking “did we validate our own generated data” would have surfaced that trust boundary before the file doubled in size.

AWS, us-east-1, 2025-10-19/20. Their own words: “a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record for the service’s regional endpoint” — two DNS Enactors raced, and cleanup deleted the stalled plan. What changed: the DNS Planner and Enactor were disabled worldwide, and the network-load-balancer change is recorded as a “velocity control mechanism to limit the capacity a single NLB can remove”. The gap: a race condition between two automated components is exactly what a postmortem’s contributing factors are built to surface after the fact — which automated actor acted on what assumption, and when.

Google Cloud, Service Control, 2025-06-12. Their own words: “a policy change was inserted into the regional Spanner tables that Service Control uses for policies…This policy data contained unintended blank fields”; it “exercised the code path that hit the null pointer causing the binaries to go into a crash loop.” What changed: Google Cloud said it would “enforce all changes to critical binaries to be feature flag protected and disabled by default” and “modularize Service Control’s architecture, so the functionality is isolated and fails open.” The gap: a blank-field input reaching production code is a data-quality failure a corrective-action register would name with a verification criterion — “a build with blank policy fields fails a pre-deploy check” — rather than a promise to test harder.

Datadog, 2023-03-08. Their own words give the outage two end times, which is this training’s whole argument about clocks: it began 06:03 UTC on 8 March across the US1, EU1, US3, US4 and US5 regions, Datadog “declared all services operational in all regions on March 9, 2023, 08:58 UTC” — about twenty-seven hours — and recorded it “fully resolved on March 10, 2023, at 06:25 UTC once we had backfilled historical data”, about forty-eight. The cause, in their words: “a security update to systemd was automatically applied… caused a latent adverse interaction in the network stack… systemd-networkd forcibly deleted the routes managed by the CNI plugin (Cilium)” — CNI being the plugin that manages a container’s network routes. Their stated lesson, worth quoting exactly: “Most important, usable live data and alerts are much more valuable than access to historical data.” What changed: the legacy update channel that applied the systemd update automatically was disabled, and the platform set out to prove it “can operate in degraded conditions without being completely down.” This is 3.3’s forty-eight-hour case, named there without the vendor — it is this one.

CrowdStrike, 2024-07-19. Blog-level evidence only, and the gap is which document you cite rather than whether one exists. The blog says “the Channel File 291 scenario is now incapable of recurring” and names staged deployment for that class of update, but it carries no mechanism a reader can check. CrowdStrike’s separate post-incident report does: “problematic content in Channel File 291 resulted in an out-of-bounds memory read triggering an exception”, which “could not be gracefully handled, resulting in a Windows operating system crash.” That is the document a record cites. Quoting the blog and calling it a root cause is the error being demonstrated here — the mechanism was published, one document over, and the research set that stopped at the blog never reached it.

Reading five providers this way turns a headline into a specific, sourced sentence a record can cite when someone later asks where a claim about a competitor’s outage came from.

One grid, two clocks: the Iberian blackout

The Iberian Peninsula blackout began 2025-04-28. ENTSO-E (the European association of electricity transmission operators that investigated it) published a factual report on 2025-10-03 and a final report on 2026-03-20 — roughly eleven months after the event. Their stated cause, in their own words: “a combination of many interacting factors, including oscillations, gaps in voltage and reactive power control, differences in voltage regulation practices, rapid output reductions and generator disconnections in Spain, and uneven stabilisation capabilities.” Their stated response: “Strengthened operational practices, improved monitoring of system behaviour and closer coordination and data exchange among power system actors.”

Set that timeline beside Cloudflare’s write-up, published within days of its own incident. A utility runs on both timescales at once: a same-week internal accounting for its own systems, and an eleven-month multi-party investigation for a cross-border grid event. A reader building Meridian’s own record should set that eleven-month expectation with leadership before the next consequential incident, not while everyone is waiting on it — a calendar-time expectation worth naming in advance, not discovering under pressure.

What your own record would have captured

Read together, the cases map cleanly only when the record’s section boundaries remain intact. Google Cloud’s change is a section 10 preventive action; Datadog’s live-data lesson is a section 2 impact finding. AWS’s automated race and Cloudflare’s trust boundary belong in the postmortem’s contributing factors and then in section 10 corrective actions. CrowdStrike’s unsupported assurance belongs in the claim registry. Section 4 remains reserved for human response decisions, and section 5 for regulatory determinations.

Stop and escalate when only a secondary summary of a provider’s incident exists and no primary write-up can be found — escalate to whoever owns your own claim registry before quoting a mechanism, a duration, or a percentage you cannot verify, because a paraphrase of a postmortem is not the postmortem, and this is a domain where paraphrases have already been shown to invent details the original never published.

KNOWLEDGE CHECK

Reading Google Cloud's June 2025 Service Control write-up, a teammate summarizes it as: 'A bad policy update caused a crash; they fixed it by testing more carefully.' What is wrong with that summary?

Key takeaway

Five corporate providers and one grid association published their own words about six selected public incidents: cause and change where stated, and an explicit evidence gap where they were not — Cloudflare’s configuration-file hardening, AWS’s race condition and velocity control, Google Cloud’s feature flags disabled by default plus a separate fail-open redesign, Datadog’s five-region outage with two end times twenty-one hours apart, CrowdStrike’s blog-level assurance beside a root-cause report the blog does not link, and ENTSO-E’s Iberian report. ENTSO-E’s report shows the other timescale: a factual report about five months out and a final report about eleven months out, beside Cloudflare’s same-week write-up. Reading a postmortem means quoting it, never paraphrasing it, and asking what your own record’s sections would have forced you to write down that theirs left out. This closes Learn — Own It picks up with rehearsing the response before the next Saturday 02:14.

LEADERSHIP DECISION require that any public incident cited in
Meridian's own materials be quoted from the
provider's primary write-up, dated, never from
a secondhand paraphrase
PRACTITIONER ACTION read each case for cause in their own words,
what changed in their own words, and what your
own record's sections would have captured that
theirs did not
SUCCESS MEASURE zero cited incident facts in Meridian's own
records traceable only to a summary - every
one dated and sourced to the provider's own
text, an audit finding avoided
Search lessons