Decide Before You Know
Last reviewed · content updated
AdvancedWhat you'll learn
~18 min- Apply stop-the-bleeding, restore-service, preserve-evidence as the commander's ordering when the cause is unknown
- Name who invokes a mitigation at 02:14 and log that call, with its rationale, in the record's decision log
- Recognize when a preservation duty inverts that ordering, and run an evidence checklist before the fastest fix
What it is: the commander’s ordering under uncertainty — stop the bleeding, restore service, preserve evidence — and the one federal delta that inverts it: a preservation duty that must be satisfied before the fastest fix runs.
What it buys: a mitigation made in minutes with a documented, defensible rationale, instead of an undocumented call a review has to reconstruct from memory — and, where a preservation clause is live, evidence that survives the fix rather than evidence the fix destroyed.
What to fund: a timestamped decision log built into the incident record itself, and a clause check — does this contract carry a preservation duty — run before the first incident it would ever apply to, not during it.
Before the detail — Artifact: the record’s decision log, one row per call, entered within minutes of the call, never reconstructed later. Status of what follows: the ordering is common baseline; the preservation inversion is binding only where the clause is in the contract.
Prompt first: log the call as you make it
It is 02:14. FieldDesk's error ratio has climbed for eleven minutes.Here is what is known [paste: the alert that fired, the last deployid and time, anything already ruled out] and here is the rollbackrunbook [paste].
Draft the FIRST row of the incident record's decision log: TIME the instant this call is made, not when the symptom started DECISION one action, taken now - roll back, rate-limit, or name the alternative - not a plan with steps WHAT WAS KNOWN the two or three facts actually in hand WHAT WAS ASSUMED the thing being bet on that is NOT yet confirmed WHO DECIDED the commander's name - never "the team"
Here is the counsel-approved preservation-duty entry for thissystem [paste]. If it says APPLIES, list the evidence checklistalready named in the approved response plan. If it says DOES NOTAPPLY, write NOT APPLICABLE. If it is missing or ambiguous, writeMISSING and route it to the clock desk and counsel; do not inferclause applicability.The prompt forces a named decider and stated assumption before the fix, and makes preservation an explicit gate.
Stop the bleeding, restore service, preserve evidence
Google IMAG gives the ordering: “stop the bleeding, restore service, and preserve evidence.” Cloud Modernization 4.3 (a separate training in this series) supplies the prepared payload — roll back before you debug — and DevSecOps 4.3 puts a bound on the same move: “roll back alone, in five minutes.” This lesson owns who invokes that prepared mitigation under uncertainty and how the call enters the decision log.
What was still missing at 1.1’s 02:14 opener was never the runbook — it was who may run it at 02:14. This lesson answers that: the commander invokes it, on their own authority, and writes down why.
Who invokes it, and what the record keeps
The paging vendor’s model gives the commander the same authority the ordering above assumes: “Identify investigation & repair actions (roll back, rate-limit services, etc) and delegate.” When the cause is not confirmed, the call itself is the Incident Commander’s (IC — the person making the mitigation call) to make — “It’s the call of the IC on how to proceed in cases where the cause is not positively known” — and the fallback when even the commander is unsure is to say so out loud: “If you are unsure, then announce publicly.”
What this training owns, and what neither of those sources teaches, is what happens to that call afterward: a timestamped row in the decision log, with a name attached. The committed NON-PRODUCTION example, INC-2026-041’s filled record, shows the shape — at 14:33, the entry reads “Roll back before finding the root cause,” what was known (“Pool exhaustion began within 60 s of the 1.14.3 rollout”), what was assumed (“That 1.14.3 was the cause and rollback was safe with no schema change”), and who decided (the commander, named). That row is 1.4’s artifact — the incident record’s decision log — filled in real time, not reconstructed at the review three days later from memory and chat scrollback.
House practice — not the paging vendor’s, not the ordering’s original source, and attributed to neither — adds three habits neither primary source names. A hypothesis ledger: every candidate cause gets a line, ranked, so the room can see what was considered and rejected instead of just what was tried. One change at a time: a single mitigating action per step, so its effect on the signal is legible — two changes at once and neither one’s contribution to the recovery can be read back out of the graph. And a timebox on the action itself: a stated instant, decided when the action is taken, at which an unconfirmed mitigation gets reverted or escalated rather than left running past whatever time it was supposed to prove itself by.
The preservation inversion
Here the commercial ordering is actively wrong, and only where a specific clause is in play. DFARS 252.204-7012 (the defense-contract clause carrying cyber-incident reporting and evidence duties) carries a 90-day duty to preserve images and packet captures — but only if that clause is present in the governing contract. Where it is, an evidence checklist runs BEFORE the fastest mitigation, not after: capture the image, capture the traffic, and only then apply the fix that would have overwritten both.
This is not a general caution to be careful — it is a specific, contract-conditioned inversion of the three-verb ordering that opened this lesson. A commander who rolls back on the DFARS-covered system before confirming whether that clause applies has made the same call this lesson just taught as correct everywhere else, and gotten it backward here. The fastest fix and the defensible fix are different actions, and this is the one place where choosing the fastest one first turns a service incident into a contract-level preservation failure.
Stop and escalate when the maintained profile is missing or ambiguous about preservation. The clock desk routes the question to counsel; the commander preserves evidence pending that decision when doing so is safe. Do not ask the assistant to interpret the contract or confirm applicability only after mitigation.
Eleven minutes into a climbing error ratio with the cause unconfirmed, the commander is deciding whether to roll back now. What does this lesson say the commander should do?
The commercial starting practice is mitigate first, root-cause later, with the commander’s authority to act on an unconfirmed hypothesis. The federal delta is a contract-specific preservation duty — DFARS 252.204-7012’s 90-day image and packet-capture requirement, live only where that clause is written into the governing contract — that runs an evidence checklist before the mitigation instead of after. The handoff artifact is that checklist, executed and timestamped, sitting beside the decision-log row it preceded. What is not equivalent: the fastest fix and the defensible fix are different actions, and a commander who treats them as the same one has picked the wrong one for a system this clause covers.
Practice status — among organizations running a mitigate-first incident response, commercial and federal
| Practice | Status | Also called |
|---|---|---|
| mitigate before confirming root cause | common baseline | fix-forward / roll back first |
| timestamped decision log with named commander and stated assumption | strong optional | incident decision journal |
| hypothesis ledger, one change at a time, timeboxed actions | reference-shop; house practice in this training | ad hoc troubleshooting discipline |
| evidence-preservation checklist run before mitigation | required where DFARS 252.204-7012 is in the contract; strong optional otherwise | forensic hold before remediation |
Scale: required | common baseline | strong optional | reference-shop (seen only at organizations that publish their own practice) | emerging
Key takeaway
Stop the bleeding, restore service, preserve evidence — in that order, with the commander invoking the mitigation on their own authority and a named decision-log row capturing what was known, what was assumed, and who decided. House practice — a hypothesis ledger, one change at a time, a timebox on the action — is this training’s addition, not a borrowed rule. The one place that ordering inverts is a contract carrying a preservation duty: there, the evidence checklist runs first. The next lesson picks up from the moment the mitigation lands: handing the commander seat to someone else without losing a beat.
LEADERSHIP DECISION require a timestamped, named decision-log row for every mitigation call, and confirm which contracts carry a preservation duty before the incident that would trigger itPRACTITIONER ACTION log time, decision, known, assumed, and decider at the moment of the call; run the evidence checklist first on any system a preservation clause coversSUCCESS MEASURE every mitigation call has a decision-log row within minutes of being made - zero reconstructed-from-memory rows found at review