Operating in Production: On-Call, Incident Command, and Reporting Clocks
Saturday 02:14: three things break at once and one person is awake. Build the human system around a running service - declare it, staff and run the response, communicate under clocks that each start at their own recorded trigger, and turn the incident into verified change instead of a slide deck.
6 modules · 23 lessons · ~6.6 hours · SRE and operations, on-call engineers, service owners, incident-response leads, and the leaders who staff and fund the rota
Chapter 8 in the Meridian sequence · 9 live so far · builds on Cloud Modernization Patterns and Modern DevSecOps Foundations and AI Assurance: System Risk and Release Decisions · continues in Guarded Automation: Agents That Run Operations
Saturday, 02:14. FieldDesk's API latency climbs after Friday's deploy — the rollback runbook exists, but nobody wrote down who may run it at this hour. The grid-analytics exchange with the federal energy-coordination system stalls, and it is not clear which reporting clock is running or whether it has even started. Monday's Azure bill will show a runaway job nobody paged on. One person is awake. This training builds the human system a running service needs before any of that happens: declare an incident and lower the threshold, decide who may staff the response and check that the coverage is actually staffed, name a commander who decides and does not fix, page on symptoms instead of noise, run every reporting obligation from its own recorded trigger, and turn the incident into an incident record — one document a regulator, a contracting officer, and next quarter's on-call can all read. Every lesson drives an AI CLI against a substrate that runs with no account: a burn-rate policy with its own unit tests, a tabletop exercise that writes the decision log, and a clock calculator that refuses to start a clock it cannot source. The thesis stays modest throughout: the capacity to run an incident is a funded, rehearsed system built before the first page, not something the person who happens to be awake improvises. No prior Meridian knowledge needed.
The Curriculum
The System Around the Service
Built before the first page
Meet the incident that runs while one person is awake, lower the threshold for declaring one, decide whether the response can actually be staffed, and name the one document that carries the record
The Page
A claim on someone's night
Read a page as a claim on a person, build the burn-rate policy on top of a given availability SLO with a unit-tested rule, treat the rota as a contract with a coverage attestation, and route cost anomalies as a page or a footnote
Command
A role, not a rank
Command is a role a named person holds: decide before you know the cause, hand the command off out loud with acknowledgment, and put the AI assistant to work in review mode only
Reporting Clocks and Communications
Each clock starts somewhere else
Every reporting obligation starts its own clock at its own recorded trigger; find which ones reach Meridian by contract flow-down, sector regulation, or its vendor, keep every status update naming the next one, and notify the regulator, the agency, and the contracting officer with content that holds up
Learn
Change something, not just morale
Run the blameless review as one document, reconstruct the timeline from logs and traces, write action items with an owner and closure evidence, and read other people's postmortems in their own words
Own It
Fund it, rehearse it, prove it
Rehearse the response with exercises that are not restore drills, cost the operations bill in people and tiers, and replay Saturday 02:14 end to end into a complete non-production incident record
New to AI CLI tools?
This training assumes you can drive an AI CLI (Claude Code, Codex CLI, Antigravity CLI, or Copilot CLI). If that's new, these modules from our AI-Powered Development training are the fastest preparation — most students need only the first one: