intellimetrics Learning
AI Assurance: System Risk and Release Decisions

Leadership brief · one page

AI Assurance: System Risk and Release Decisions

Most organizations now run more AI systems than they can list, and none of them carries a signature. This training builds the release decision as a capability - inventory, consequence tier, evidence, evaluation on your own terms, a signed and expiring record - and is honest that the record is a bounded claim, never a certification.

Self-paced · 6 modules, 23 lessons · about 7 hours

Chapter 7 in the Meridian sequence · 9 live so far — the sequence follows one fictional utility through the same modernization, so the examples build on each other, and this one picks up the story from Grounded Answers From Documents and Zero Trust Implementation. Each training stands on its own; the order is the recommended path, not a prerequisite — the examples name the same fictional utility, but nothing from the earlier trainings is needed to follow them. The story continues in Operating in Production: On-Call, Incident Command, and Reporting Clocks.

What this training covers

Six modules in the order the work actually happens: listing every AI system with an owner and tiering each by consequence, including the exit where a system is not deployed because its gates will not be funded; assembling what the model and its provider actually claim, and the version you can actually pin; evaluating on a sealed set the provider never saw, differently for a classifier, a generator, and a system you cannot open; challenging that evidence by attempt budget; putting the decision on a signed record with an independent reviewer, verifiable conditions, and the determinations and waivers that must be written down; and keeping the decision current until the day you withdraw it. Every module teaches the work and the AI-assisted way to do it as one workflow, and every module ends by finding what the agent got wrong.

Why it matters

The discipline most organizations borrowed for validating models - banking model-risk management - now says it does not cover this: the revised US guidance of April 2026 (Federal Reserve SR 26-2 and OCC Bulletin 2026-13) states that generative and agentic AI are outside its scope, and OMB's consolidated 2025 federal AI use case inventory (the count in its published inventory README) lists 3,611 individually reported use cases, 445 of them high-impact - each one somebody had to sign for. The defensible version of this capability makes a modest claim: a release decision is a bounded, signed, time-limited statement about a named system and a named version, supported by dated evidence and reviewed by someone who did not build it - never, by itself, a certification or an authorization to operate. Everything in this training exists to keep the decision inside that claim: a register you can trust, gates chosen by consequence, evidence you produced on your own terms, and an expiry that makes the provider's retirement notice a trigger rather than a surprise.

What changes in practice

Tags name what each shift affects most: calendar time, cost, contract risk, or an audit finding avoided.

  1. 1

    Every AI system is on a register with a named owner and a version boundary

    Audit finding avoided

    Why it matters: you cannot assure what you cannot list - and the register is where the board question "which AI do we run, and who signed for it" gets an answer in one minute instead of a quarter · Module 1

  2. 2

    Consequence decides the gates, and one exit is "do not deploy"

    Cost

    Why it matters: high-impact means the output is the principal basis for a decision with legal, material, or safety effect on a person - a test, not an owner's election - and a high-impact system whose gates nobody will fund stops before it ships, which is cheaper than stopping it after · Module 1

  3. 3

    The vendor's card is recorded as a claim; the version you can actually pin is written down

    Contract risk

    Why it matters: model names hot-swap and retire on the provider's clock, sometimes within weeks - the record names the immutable version, the notice date, and the alias you could not pin · Module 2

  4. 4

    Evaluation runs on a sealed set the provider never saw, with separate gates per system type

    Cost

    Why it matters: a classifier, a generator, and a system you cannot open fail in different ways - and a set the vendor tuned against measures the tuning, not the system · Module 3

  5. 5

    Attack evidence is reported by attempt budget, never by one lucky run

    Contract risk

    Why it matters: a single-attempt success rate is insufficient wherever retries are plausible - the curve at one, ten, and a hundred attempts, with its uncertainty, is what the decision rests on · Module 4

  6. 6

    The decision is a signed record with an independent reviewer, verifiable conditions, and an expiry

    Audit finding avoided

    Why it matters: the person accepting the risk signs; someone who did not build or buy the system records the findings; every condition has an owner and an enforcement point - and the record says only "passed the named gates on the named set" · Module 5

  7. 7

    A fired trigger starts a clock, and an unanswered clock expires the decision

    Calendar time

    Why it matters: the provider's retirement notice, a crossed version boundary, or a changed context suspends or expires yesterday's decision - and the portfolio view shows what expires this quarter before it does · Module 6

Where the effort goes

Most of it lands in the register, the impact assessment, and the reviewer's hours, not in tooling. Naming an owner and a version boundary for every system, writing down the current practice a system replaces and what it costs when it fails, and having someone independent read the evidence is the real work; the evaluation harness runs in minutes once an agent is driving it. What this does not require is a governance-platform purchase or a new model contract - the practice substrate runs on your existing AI CLI plus a small local model and a mock, with no account and no admin request. What it does require is funded time: assessor-hours per tier, a reviewer who did not build the system, and the discipline to let "do not deploy" and "withdraw" stand as ordinary outcomes.

How you'll know it worked

If you read one lesson, read The Release Decision (Lesson 1.4). It turns the bounded claim above into a record someone can sign, and names who signs and who reviews before any test runs. It's written for you, not just for your engineers.

If the work lands on you, start at the curriculum page — 6 modules in dependency order, opening with Ten Systems, One Signature (Lesson 1.1). Inside this training the modules are sequential - each depends only on what came before.

Every lesson ends with the same three lines: the decision that is yours, the action that is your team's, and the measure that says it worked. If someone sends you a lesson, read those three lines first.