Operating in Production: On-Call, Incident Command, and Reporting Clocks Module 1 · The System Around the Service

Saturday 02:14

Last reviewed · content updated

Beginner

What you'll learn

~12 min
  • State the training's thesis and name the decision each of its five corollaries governs
  • Identify the three roles a running incident needs before the first page, not during it
  • Draft, from an org chart and its contracts, who would be paged, who may declare, and which clocks could start

Before the detail — Decision: treat an incident as a decision problem with named owners, not a technical problem for whoever is awake. Outcome: three unresolved threads from one bad night, and the one artifact this training leaves behind to close them cleanly next time. Artifact: the incident record (Lesson 1.4). Status of what follows: reusable guidance — this lesson sets the frame the rest of the training makes binding.

Prompt first: find out who is actually named

Here is our team's org chart [paste] and the customer or partner
contracts that mention incident response, notification, or uptime
[paste relevant clauses only].
For each of these three questions, give me a named person or role,
or write MISSING if the org chart and contracts do not answer it:
1. WHO WOULD BE PAGED first if our main service went down at 02:14
on a Saturday?
2. WHO MAY DECLARE an incident and choose its severity - does that
authority sit with a role, or only with whoever happens to be on
shift?
3. WHICH CLOCKS could start from an event like this - a contract
clause, a regulator, a customer commitment - and who is named to
decide whether one has?
Do not guess a name to fill a gap. MISSING is the correct answer
where the org chart has no answer, and it is more useful than a
guess.

The assistant can only read what you paste, so the prompt asks it to name gaps instead of papering over them. A chart with three MISSINGs is worth more on a Tuesday afternoon than a confident guess is worth at 02:14.

Three failures, one person awake

It is 02:14 on a Saturday. J. Whitfield, Meridian’s on-call platform engineer, is the only person looking at a screen, and three things are wrong at once.

FieldDesk’s API latency has been climbing since Friday’s deploy. Cloud Modernization 4.3 — a separate training in this series — has the rule for this: roll back before you debug; its runbook exists. But nobody has written who may run it at 02:14, on whose authority, with the field engineering manager asleep and R. Okafor, FieldDesk’s service owner, not yet paged.

Second, the grid-analytics exchange with the federal energy-coordination system has stalled. Something may have just started a clock — a determination, a discovery, an obligation with a deadline attached — and Whitfield has no way to know which, or whether it has, or who is supposed to decide.

Third, and nobody will notice until Monday: a job somewhere in Azure has been running since Friday night with nothing watching its cost, and the bill will say so before any dashboard does.

One engineer, three threads, and none of them were designed to be solved by the one person who happens to be awake — every hour spent deciding who is even allowed to act is an hour the contract, the exchange, and the bill are all left unattended at once.

The thesis

An incident is a decision problem run by named people: they declare, command, communicate under each applicable clock from its own recorded trigger, preserve evidence, and turn one incident record into verified change — and the capacity to do that is a funded, rehearsed system that exists before the first page, not something the person who happens to be awake improvises.

Five corollaries follow, one per group of chapters ahead:

  • Every page is a claim on a person’s night. An unactionable page is a defect, not noise — Whitfield’s pager should never fire for something nobody can act on.
  • Declare low, by consequence-to-whom. A severity ladder measures the response you can mount; a federal major-incident determination measures harm on criteria you do not control. Your SEV-1 is not a major incident.
  • Command is a role, not a rank. The commander decides and does not fix; handoff is spoken and acknowledged; when the cause is unknown, the order is stop the bleeding, restore service, preserve evidence — and where a clause carries a preservation duty, that order can invert.
  • Each obligation starts at its own recorded trigger. Determination, discovery, and materiality diverge, every update names the next one, and publishing never discharges a duty to notify.
  • The review’s product is change, not a scorecard. MTTR is not a per-incident scorecard; an assistant may draft while a person decides; and a fuller kind of autonomy is a door a later chapter opens, not this one.

Whitfield’s three threads sit inside these five: the rollback is corollary three’s territory, the stalled exchange is corollary four’s, and the unwatched job is a page that should have fired under corollary one.

Why this stopped being a technical problem

Two facts explain why an unstaffed on-call rotation is a leadership problem. The National Institute of Standards and Technology’s April 2025 incident-handling guidance, Special Publication 800-61 Revision 3 (SP 800-61r3), contrasts incidents once “completed within a day or two” with recovery that “takes weeks or months.” The SRE Report 2026 (Catchpoint, n=418, published 2026-01-22) found that only 22% of respondents said their organizations financially model the cost of downtime or degradation in ways that inform decisions; that finding describes the surveyed respondents, not the whole industry — Whitfield’s three threads are still costing something nobody has a number for before Monday’s meeting.

Neither fact is a reason to panic. Both are reasons to fund a system before the next 02:14.

One artifact holds the whole answer

Everything this training builds — declaring, commanding, running the clocks, communicating, learning — writes into one document: the incident record. It is not a ticket and not a chat log. It has a state, a timeline, decisions with their rationale, dated determinations, the clocks that ran, and a register of what changed afterward, and a regulator, a contracting officer, and next quarter’s on-call engineer can all read the same pages. Lesson 1.4 opens it section by section; this lesson only names it, because naming the destination is what keeps the next three lessons from feeling like separate problems.

Stop and escalate when the prompt above comes back MISSING for any of the three questions — who would be paged, who may declare, which clocks could start. A named gap in an on-call chart is not a problem this training expects you to quietly patch; it goes to whoever owns coverage decisions for that service, in writing, before the next real page — not during it. Lesson 1.3 is where that decision gets made.

KNOWLEDGE CHECK

Whitfield already has Cloud Modernization 4.3's rollback runbook, and could technically run it at 02:14. According to this lesson's thesis, what is actually missing?

Key takeaway

An incident is a decision problem, not a technical one solved by whoever is awake — the capacity to declare, command, communicate under a clock, and turn one incident record into verified change has to be designed, funded, and rehearsed before the first page. Whitfield’s three threads at 02:14 map onto the training’s five corollaries and six modules ahead. The next lesson gives the first of those decisions a name: what actually counts as declaring an incident, and how low that threshold should sit.

LEADERSHIP DECISION fund an on-call system before the next 02:14,
not after it - this training's five corollaries
name what "funded" has to cover
PRACTITIONER ACTION run the prompt above against your own org chart
and contracts; take every MISSING to the person
who owns coverage decisions for that service
SUCCESS MEASURE zero MISSING answers for who is paged, who may
declare, and who owns a clock - a gap closed in
calendar time now, not discovered during a real
incident
Search lessons