Operating in Production: On-Call, Incident Command, and Reporting Clocks Module 1 · The System Around the Service

Coverage Is a Decision

Last reviewed · content updated

Advanced

What you'll learn

~18 min
  • Apply on-call capacity inputs (incident load, the operational-time budget, backup and spacing guidance) to size a rota that a real shift can carry
  • Distinguish a health cap that limits a person from a contractual floor the organization must staff
  • Route a coverage question to one of four exits, including the one where the commitment does not get signed
ℹLeadership brief

What it is: the coverage decision — a written test, before any commitment is signed, of whether the organization can actually staff the on-call response its contracts and its own severity ladder require.

What it buys: a documented stop before a 24×7 promise gets made on a rota that cannot carry it, instead of a contract breach discovered the first weekend nobody answers.

What to fund: enough named people, at enough tiers, with a real backup schedule — or the decision to contract the work out, renegotiate the commitment in writing, or not make the commitment at all.

Before the detail — Artifact: a coverage attestation naming people per tier, tiers, backups, and the contractual floor it meets. Status of what follows: binding wherever a signed 24×7 commitment exists; common baseline everywhere else.

Prompt first: size the rota against what a shift can carry

Here are our severity matrix, dated on-call schedule (window,
primary, backup, and tier), compensation-cap policy, spacing
guidance, and exact coverage clause [paste; use NONE if absent].
For each person, calculate incidents per 12-hour shift, total
operational-work percentage, on-call percentage, backup coverage,
and spacing. Name every failed constraint and show the smallest
roster that meets all constraints at once.
Do not recommend hiring and do not invent a constraint; show only
the arithmetic from the supplied artifacts.

The agent is useful here as a calculator against numbers a human supplies, not as a source of the numbers themselves — the caps above come from published on-call guidance and the reader’s own artifacts, and the one thing worth checking by hand afterward is whether the “smallest roster” answer actually meets every constraint at once, not just the last one applied.

The coverage decision, not a schedule

A rota is a schedule. Coverage is a decision: given what the organization has promised — in its own severity ladder, and in any contract that names an uptime or a response window — can it actually staff the response, tier by tier, week after week, without burning out the people who carry it? That question gets answered once, in writing, before a commitment is signed, not discovered the first weekend nobody picks up.

This training calls it the coverage decision, deliberately not borrowing another training’s name for a different kind of gate, because it produces a written attestation, not a pass/fail stamp on a system.

What a shift can actually carry

Google’s public engineering guidance on running on-call rotations sets a capacity limit in incidents, not pages: “the maximum number of incidents per day is 2 per 12-hour on-call shift.” A rota that pages a person into three or more genuine incidents a shift is not understaffed by a little — it is asking one person to run more decision problems per night than the guidance says a person can carry well.

The same source sets one operational-work budget: at most 50% of total time on purely operational work, with no more than 25% on-call and up to another 25% on other operational, nonproject work. These are parts of one ceiling, not independent allowances.

Two more published guidance inputs remain, both dated vendor recommendations rather than Meridian mandates. One widely used paging vendor’s on-call guide is blunt about backups: “Always have a backup schedule. Yes, this means two people being on-call at the same time,” with an escalation to the next tier after five minutes if nobody acknowledges. And guidance updated 2026-04-27 from a separate incident-management vendor recommends at least two weeks between one person’s shifts, ideally one week on and three weeks off. Using that ideal pattern as the scenario assumption yields four people per tier; that is a derived example, not a rule.

None of these four inputs is a mandate on Meridian. Each is attributed guidance — a vendor’s dated recommendation, a public engineering practice — and the coverage decision is what Meridian does with them: run the numbers against the real roster, and see what falls out.

A health cap limits a person; a contract floor is something you must staff

The four inputs above are all health caps: they describe what keeps one person’s on-call load sustainable, and where a rota cannot meet them, Meridian can choose to accept the risk, hire, or change the severity ladder that is driving the load.

A contractual floor is a different kind of number. Where Meridian’s contract with the state-federal interconnect commits to 24×7 coverage of the grid-analytics exchange, that commitment is not a target to be balanced against a health cap — it is a floor the organization must staff, in writing, or renegotiate in writing, or decline to sign. A health cap that is not met is a risk somebody can choose to accept. A contract floor that is not met is a breach somebody signed up for without the staff to back it.

Conflating the two is the mistake this lesson exists to prevent: a risk acceptor who is comfortable accepting a health-cap gap is not thereby authorized to accept a contractual one, because the customer on the other end of that contract never agreed to accept it too.

Four exits, one signature

Every coverage question resolves through exactly one of four exits, and all four are legitimate outcomes — including the last one.

Staff it. A funded rota that meets the health caps and the contractual floor, named person by named person, tier by tier, with backups. This is the outcome the arithmetic above is meant to produce when the roster can carry it.

Contract it. An outside provider staffs the gap under a written handoff that names what they own and what Meridian still owns. The buy itself — vetting and contracting a managed operations provider — is Software Acquisition’s territory, not this training’s; this lesson only names contracting as a legitimate exit from the coverage decision.

Obtain a written contractual modification. If the floor as signed cannot be staffed and cannot be bought, the organization goes back to whoever holds contract authority and changes the commitment in writing — a shorter coverage window, a different response time — before the next incident tests the gap that was never closed.

Stop: do not sign a coverage commitment you cannot staff. This is the cheapest exit in the whole training — the risk is removed for the price of a signature that was never given — and it is available right up until someone signs.

A named risk acceptor — Meridian’s P. Delgado, in this training’s cast — can accept degraded coverage in writing only where no binding floor exists. Delgado cannot waive a customer’s contractual right to 24×7 coverage on the organization’s behalf; that authority sits with whoever holds contract authority, not with the person accepting operational risk. Data Products 6.1, a separate training in this series, asks who is woken at 4am; this is where that question stops being a hypothetical for leadership and becomes a signature.

Stop and escalate when a coverage attestation would require the risk acceptor to waive a contractual 24×7 floor in order to sign it. That is not a risk-acceptance decision — it goes to whoever holds contract authority, in writing, before the shift in question starts, not after a page goes unanswered.

KNOWLEDGE CHECK

Meridian's FieldDesk rota meets every health cap - incident load, the operational-time budget, backup coverage, and shift spacing - but the state-federal interconnect contract commits to a 24×7 response window the current roster cannot actually cover during the third week of every month. Delgado, the risk acceptor, is comfortable accepting the gap in writing. Is that sufficient?

Key takeaway

Coverage is a decision made once, in writing, before a commitment is signed: run the incident-load cap, the operational-time budget, backup coverage, and shift spacing against the real roster, and separately check whether a contractual floor exists that no health cap can override. Four exits follow — staff it, contract it, obtain a written modification, or stop — and a risk acceptor can only sign for the ones a contract has not already spoken for. Lesson 1.4 turns to where all of this gets written down for good: the incident record itself.

LEADERSHIP DECISION fund the rota the coverage decision calls for,
sign a written contract modification, or accept
the stop exit - the one option not available is
a 24×7 promise nobody can actually staff
PRACTITIONER ACTION run the health-cap arithmetic against the real
roster, separate any contractual floor from it,
and route the answer to one of the four exits
with a named signature
SUCCESS MEASURE a written coverage attestation for every
service under a 24×7 commitment - zero unsigned
gaps discovered live - contract and revenue
risk removed before the next shift, not after it
Search lessons