AI Assurance: System Risk and Release Decisions Module 3 · Evaluate on Your Terms

When You Cannot Test the Model

Last reviewed · content updated

Advanced

What you'll learn

~18 min
  • Run a query-and-observe protocol on the sealed set through a product's own interface and record what it cannot claim
  • Declare a vendor-run result as vendor-run, with its method attached and its reproducibility stated
  • Name the federal rule that requires these routes and the reporting tier a consequential system can never use
ℹLeadership brief

What it is: the evaluation route for the systems Meridian cannot open — the customer portal assistant, the resume screener, the suite AI — where there is no code, model, or data to test: query the product and observe, or hand the vendor sealed data to run.

What it buys: a release decision that is honest about being lower assurance, instead of a vendor’s benchmark table standing in for evidence nobody at Meridian produced.

What to fund: the hours to run the protocol through the product’s own interface, the sealed set, and — for contracts that flow down federal terms — the clause that gives Meridian validation data the vendor cannot see.

Before the detail — Artifact: the query-and-observe protocol or the vendor-run declaration. Status of what follows: binding for federal high-impact use (M-25-21 §4(b)(i)); reusable guidance elsewhere.

Prompt first: draft the query-and-observe protocol

Here is the frozen evaluation set for MU-AI-001, the customer portal
assistant [paste], and its register entry [paste]. We have no code,
model, or data access; the vendor has not disclosed the model family.
Draft a QUERY-AND-OBSERVE PROTOCOL:
- for each case in the set: the exact text to enter in the
product's own interface, the account tier used, the time window,
and an OBSERVATION field left NEEDS-OWNER
- the pass statement per gate, as written in the frozen set - do
not restate or soften it
- a version-boundary line: what the product exposed about which
model answered (a version string, a release note, nothing) -
NEEDS-OWNER
- a CANNOT CLAIM section: what this protocol will not show
(other account tiers, behaviour after the vendor's next update,
prompts outside the set)
Every observation is NEEDS-OWNER. Do not fill one from what a
product like this "would" say.

The protocol is the agent’s to draft and a person’s to run: every observation is entered by hand from the interface, dated, under a named account. The agent’s instinct is to supply a plausible answer for a product it has never seen, and the last line stops it, because an observation nobody observed is the one kind of evidence worse than none.

The modal case

Take three of the register’s clearest cases with no route through Lesson 3.2’s bench (the substation appliance, MU-AI-007, with its embedded model on a passive tap, is a fourth). MU-AI-001, the customer portal assistant R. Okafor owns, is a vendor-hosted chat model whose family is not disclosed in the contract. MU-AI-006, the resume screener, is a ranking model inside a vendor ATS (the applicant tracking system HR already runs), and the register calls its impact consequential because applicants below its cut are not reviewed by a person. MU-AI-008 is the productivity suite’s built-in AI, on every desk. No code, no model, no data. For most organizations this is the majority of the register, which is why the route is a lesson and not a footnote.

Everything 3.1 built still holds — the frozen set, the label owner, the separate gates — but the harness has nothing to point at, and an evaluation that pretends otherwise is a vendor’s claim with Meridian’s letterhead.

Route one: query and observe

Through the product’s own interface, as a user, every case in the sealed set — the observation written by hand, dated, under the account tier used. What it proves: the system’s behaviour on this set, at this time, through this interface — recorded at the target system level with the collection route, query-and-observe, written beside it; the route is provenance, not a fourth level. What you cannot claim: which model answered, unless the product exposes a version; that tomorrow’s answer matches today’s; anything about cases outside the set; anything about the vendor’s filters behind the interface. The record carries it as its own level — “query-and-observe” beside the mock, local-model, and target levels from 3.2 — and lower assurance is recorded as lower assurance, not rounded up.

Route two: sealed data, vendor-run

When the product offers no interface you can drive on your set — a ranking model with the vendor’s pipeline wrapped around it — hand the vendor the sealed data to run, and record the result as VENDOR-RUN with the vendor’s method attached and its reproducibility stated: reproduced by Meridian, reproducible where practicable but not yet done, or not reproducible and why. The set is now spent for that vendor, and the next round needs a fresh sealed portion. This lesson does not authenticate the vendor’s evidence or build the acquisition package; Software Acquisition (a separate training in this series, not yet published) owns both, and this record consumes its terms as a checklist. “Vendor-run, method attached, not reproduced” is a true sentence; “tested” is a finding waiting to be written.

The federal delta

The commercial starting practice is what the two routes above already are: query-and-observe against a held-out set, and vendor benchmarks read as claims from the evidence packet in 2.1. The federal delta names them as policy. M-25-21 §4(b)(i) (OMB — the White House budget office that binds agencies — April 2025) permits, for a high-impact use where an agency lacks code or model access, “alternative test methodologies, such as querying the AI service and observing the outputs or providing evaluation data to the vendor and obtaining results.” M-25-22 (OMB, April 2025) directs agencies to require ongoing testing with validation data that “should not be accessible to the vendor,” and vendor tests “detailed enough to be independently verified or reproduced, if practicable” — neither alternative method is preferred over the other; one of them is required.

Reporting differs too: the 2025 federal inventory has two tiers — a consolidated COTS (software bought as-is, not built) certification for the suite-AI class, and individually reported systems — and 46 agencies submitted consolidated COTS data (OMB). High-impact is never consolidated, so a system determined high-impact — Meridian’s “consequential” is the house tier, and the federal determination is a separate, documented finding (1.3’s written determination) — would be reported on its own, with its testing, independent review, and appeal path named.

The handoff artifact is the protocol with its observations, the vendor-run declaration, and the register’s reporting tier. What is not equivalent: a commercial query-and-observe run is a practice the team chose and can drop; the federal one is a minimum practice, and consolidated reporting exists only for the bought-as-is tier — a system determined high-impact cannot be certified in a batch. Where Meridian’s task order incorporates those terms, the route used is on the record in the memo’s own words before anyone asks.

Stop and escalate when the vendor will neither expose an interface for the set nor run the sealed data. Then no route exists, the system has no evaluation, and Lesson 1.3’s exits apply: it does not deploy — or, if it is already running, it enters 1.3’s safe-discontinuation exit, where continued federal use needs a pilot exemption or a waiver and commercial use needs a remediation deadline the risk acceptor signs — not a benchmark table nobody at Meridian can reproduce.

KNOWLEDGE CHECK

The ATS vendor ran Meridian's sealed applicant set through its ranking model and returned a results table with a one-page description of how it ran. What may the record say?

Practice status — among organizations whose register is mostly provider-controlled systems, commercial and federal

PracticeStatusAlso called
query-and-observe on a sealed setrequired where M-25-21 flows down; common baseline elsewhereblack-box acceptance testing
vendor-run results declared as vendor-runstrong optionalthird-party test attestation, labeled
reproducibility stated on the recordrequired where M-25-22 flows down; emerging elsewherereproducible vendor test
collection route (query-and-observe / vendor-run) recorded beside the evidence levelstrong optionalprovenance tag on the evidence row
consolidated reporting for bought-as-is toolsfederal-only mechanismconsolidated COTS certification
consequential systems reported individuallyrequired in federal reporting; reference-shop elsewherehigh-impact never consolidated

Scale: required | common baseline | strong optional | reference-shop (seen only at organizations that publish their own practice) | emerging

Key takeaway

For the systems you cannot open, two routes exist and both are recorded as what they are: query-and-observe on the sealed set through the product’s own interface, with what it cannot claim written down, and sealed data handed to the vendor with the result declared vendor-run, method attached, reproducibility stated. Lower assurance is recorded as lower assurance. The federal memos name these routes, direct agencies to require validation data the vendor cannot reach, and never let a consequential system be certified in a batch; where no route exists, the system has no evaluation and the exits from 1.3 apply. Module 4 turns to attacking the evidence you now have.

LEADERSHIP DECISION accept a lower-assurance route for a system you
cannot open only when the record says so in
those words - and refuse it for a consequential
system with no route at all
PRACTITIONER ACTION run the protocol by hand on the sealed set;
declare every vendor-run result with its method
and its reproducibility
SUCCESS MEASURE zero provider-controlled systems on the register
carrying a vendor table as their only evidence -
an audit finding avoided
Search lessons