Leadership brief · one page
Operating in Production: On-Call, Incident Command, and Reporting Clocks
Most organizations can name their software; few can name who is on call for it, which clock started when it broke, and what changed afterward. This training builds the incident as a run system - declare it, put a named commander in charge, communicate under every clock it starts, and turn the record into verified change - and is honest that a page is a claim on a person's night, not a suggestion.
Self-paced · 6 modules, 23 lessons · about 7 hours
Chapter 8 in the Meridian sequence · 9 live so far — the sequence follows one fictional utility through the same modernization, so the examples build on each other, and this one picks up the story from Cloud Modernization Patterns, Modern DevSecOps Foundations, AI Assurance: System Risk and Release Decisions. Each training stands on its own; the order is the recommended path, not a prerequisite — the examples name the same fictional utility, but nothing from the earlier trainings is needed to follow them. The story continues in Guarded Automation: Agents That Run Operations.
What this training covers
6 modules in the order an incident actually runs: declaring it, deciding whether the response can be staffed, and naming the one record that carries the timeline forward; treating a page as a claim on a person and building the burn-rate policy that decides which pages are real; putting a named commander in charge who decides and does not fix, with an artificial intelligence (AI) assistant that drafts under review and never acts alone; running every reporting obligation - to a regulator, a sector regulator, or through Meridian's own vendor - from its own recorded trigger, never a single shared deadline; turning the review into corrective actions with an owner, a due date, a verification criterion, and closure evidence, instead of a metric nobody can act on; and rehearsing, staffing, and costing the whole system before the next incident finds the gap. Every module teaches the work and the AI-assisted way to do it as one workflow, using a practice substrate that runs with no account.
Why it matters
Detection tooling is only one part of the human system that has to act on what gets detected. NIST (the federal standards body for cybersecurity guidance) published SP 800-61r3 in April 2025 as incident-handling guidance, not a mandate; it says recovery can take "weeks or months" and that lessons learned "should often be shared as soon as they are identified, not delayed until after recovery concludes" — exactly the gap a funded, rehearsed system closes. The 2026 Site Reliability Engineering (SRE — operating services against explicit reliability targets) Report (n=418 respondents, published 2026-01-22) found that only 22% financially model the cost of their own downtime, which means most on-call decisions get made without the one number that would justify funding them. And the clock that decides whether a report is late is not the one most leaders assume: a memo from OMB (the White House office that sets federal reporting policy), M-25-04, footnote 42, starts the one-hour federal incident-reporting clock at the moment an agency determines an incident is major - not at the moment it was first detected. Meridian is a vendor, not a federal agency, so obligations like that one reach it three ways: through a contract's flow-down clause, through the sector regulator that already licenses it, or through its own vendor's federal authorization - never by filing with Congress directly. Lesson 4.2 walks the three paths so a reader can tell which one a given contract, regulator, or vendor puts them on. A coverage decision, a named commander, and a determination entry recorded on time are what keep an organization compliant and defensible on the same night.
What changes in practice
Tags name what each shift affects most: calendar time, cost, contract risk, or an audit finding avoided.
- 1
Coverage is decided, staffed, and put in writing before a page ever fires
Contract riskWhy it matters: a person cannot be paged into existence at 2am - the coverage decision names who staffs the response, and a risk acceptor (the named person authorized to accept reduced coverage for the organization) can accept degraded coverage only where no binding contract floor already exists · Module 1
- 2
Every page passes a fatigue-and-symptom test before it reaches a phone
CostWhy it matters: an unactionable page is a defect, not noise - a burn-rate policy with its own passing test decides what actually wakes someone, so on-call capacity goes to real fires instead of a pager that cried wolf · Module 2
- 3
Command is a named role, and the commander decides without touching a keyboard
Calendar timeWhy it matters: one source of truth for what is happening and what happens next reduces response delay, and a spoken, acknowledged handoff keeps command continuous past any one shift · Module 3
- 4
Every reporting obligation runs its own clock from its own recorded trigger
Audit finding avoidedWhy it matters: Meridian can be compliant and late on the same event if the wrong clock is watched - a determination entry with a timestamp and a named decider is what proves the right clock started on time · Module 4
- 5
The review changes something, and one document serves the team and the record
Audit finding avoidedWhy it matters: a blameless internal review and an external investigation are not the same document - corrective actions with an owner, a due date, a verification criterion, and closure evidence are what turn a lesson into a change instead of a slide · Module 5
- 6
The operations bill is staffed, rehearsed, and costed before the next incident, not after
CostWhy it matters: exercises that are not restore drills, a costed coverage attestation, and a status page are cheaper than discovering the gap live — NERC CIP-008 (the grid cyber-security standard) tests the response plan every 15 calendar months where it applies, while M-17-12 (the federal breach-response memo) requires covered agencies to run a tabletop not less than annually · Module 6
Where the effort goes
Most of the effort lands in staffing and rehearsal. The work is naming who covers each shift, confirming that coverage meets a contract's floor, putting a commander in charge who stops resolving while holding the role, and running a tabletop that produces a decision log. The practice substrate uses the command-line assistant the team already has plus small tools that require no account; it does not require a new monitoring platform or incident-management contract. The recurring work is rotation seats under the compensation cap, a commander pool trained to decide without fixing, scheduled exercises, and a written modification or stop decision whenever available coverage cannot meet a contractual floor.
How you'll know it worked
- Every service has a written coverage attestation - people per tier, backups, and the contract floor it meets - signed before an incident, not discovered during one
- Every page that fires passes a symptom-and-fatigue test, and the alert library is reviewed as page, sub-critical, or delete on a real cadence
- Every incident produces one incident record a regulator, a contracting officer, and next quarter's on-call can all read, with determination entries that name a timestamp and a decider
- Every review closes with corrective actions that have an owner, a due date, a verification criterion, and closure evidence - and the next tabletop exercise rehearses the gap the last one found
If you read one lesson, read The Incident Record (Lesson 1.4). It shows the one document a regulator, a contracting officer, and next quarter's on-call can all read - and names what was AI-drafted and who verified it. It's written for you, not just for your engineers.
If the work lands on you, start at the curriculum page — 6 modules in dependency order, opening with Saturday 02:14 (Lesson 1.1). Inside this training the modules are sequential - each depends only on what came before.
Every lesson ends with the same three lines: the decision that is yours, the action that is your team's, and the measure that says it worked. If someone sends you a lesson, read those three lines first.