When You Cannot Test the Model
Last reviewed · content updated
AdvancedWhat you'll learn
~18 min- Run a query-and-observe protocol on the sealed set through a product's own interface and record what it cannot claim
- Declare a vendor-run result as vendor-run, with its method attached and its reproducibility stated
- Name the federal rule that requires these routes and the reporting tier a consequential system can never use
What it is: the evaluation route for the systems Meridian cannot open — the customer portal assistant, the resume screener, the suite AI — where there is no code, model, or data to test: query the product and observe, or hand the vendor sealed data to run.
What it buys: a release decision that is honest about being lower assurance, instead of a vendor’s benchmark table standing in for evidence nobody at Meridian produced.
What to fund: the hours to run the protocol through the product’s own interface, the sealed set, and — for contracts that flow down federal terms — the clause that gives Meridian validation data the vendor cannot see.
Before the detail — Artifact: the query-and-observe protocol or the vendor-run declaration. Status of what follows: binding for federal high-impact use (M-25-21 §4(b)(i)); reusable guidance elsewhere.
Prompt first: draft the query-and-observe protocol
Here is the frozen evaluation set for MU-AI-001, the customer portalassistant [paste], and its register entry [paste]. We have no code,model, or data access; the vendor has not disclosed the model family.
Draft a QUERY-AND-OBSERVE PROTOCOL: - for each case in the set: the exact text to enter in the product's own interface, the account tier used, the time window, and an OBSERVATION field left NEEDS-OWNER - the pass statement per gate, as written in the frozen set - do not restate or soften it - a version-boundary line: what the product exposed about which model answered (a version string, a release note, nothing) - NEEDS-OWNER - a CANNOT CLAIM section: what this protocol will not show (other account tiers, behaviour after the vendor's next update, prompts outside the set)Every observation is NEEDS-OWNER. Do not fill one from what aproduct like this "would" say.The protocol is the agent’s to draft and a person’s to run: every observation is entered by hand from the interface, dated, under a named account. The agent’s instinct is to supply a plausible answer for a product it has never seen, and the last line stops it, because an observation nobody observed is the one kind of evidence worse than none.
The modal case
Take three of the register’s clearest cases with no route through Lesson 3.2’s bench (the substation appliance, MU-AI-007, with its embedded model on a passive tap, is a fourth). MU-AI-001, the customer portal assistant R. Okafor owns, is a vendor-hosted chat model whose family is not disclosed in the contract. MU-AI-006, the resume screener, is a ranking model inside a vendor ATS (the applicant tracking system HR already runs), and the register calls its impact consequential because applicants below its cut are not reviewed by a person. MU-AI-008 is the productivity suite’s built-in AI, on every desk. No code, no model, no data. For most organizations this is the majority of the register, which is why the route is a lesson and not a footnote.
Everything 3.1 built still holds — the frozen set, the label owner, the separate gates — but the harness has nothing to point at, and an evaluation that pretends otherwise is a vendor’s claim with Meridian’s letterhead.
Route one: query and observe
Through the product’s own interface, as a user, every case in the sealed set — the observation written by hand, dated, under the account tier used. What it proves: the system’s behaviour on this set, at this time, through this interface — recorded at the target system level with the collection route, query-and-observe, written beside it; the route is provenance, not a fourth level. What you cannot claim: which model answered, unless the product exposes a version; that tomorrow’s answer matches today’s; anything about cases outside the set; anything about the vendor’s filters behind the interface. The record carries it as its own level — “query-and-observe” beside the mock, local-model, and target levels from 3.2 — and lower assurance is recorded as lower assurance, not rounded up.
Route two: sealed data, vendor-run
When the product offers no interface you can drive on your set — a ranking model with the vendor’s pipeline wrapped around it — hand the vendor the sealed data to run, and record the result as VENDOR-RUN with the vendor’s method attached and its reproducibility stated: reproduced by Meridian, reproducible where practicable but not yet done, or not reproducible and why. The set is now spent for that vendor, and the next round needs a fresh sealed portion. This lesson does not authenticate the vendor’s evidence or build the acquisition package; Software Acquisition (a separate training in this series, not yet published) owns both, and this record consumes its terms as a checklist. “Vendor-run, method attached, not reproduced” is a true sentence; “tested” is a finding waiting to be written.
The federal delta
The commercial starting practice is what the two routes above already are: query-and-observe against a held-out set, and vendor benchmarks read as claims from the evidence packet in 2.1. The federal delta names them as policy. M-25-21 §4(b)(i) (OMB — the White House budget office that binds agencies — April 2025) permits, for a high-impact use where an agency lacks code or model access, “alternative test methodologies, such as querying the AI service and observing the outputs or providing evaluation data to the vendor and obtaining results.” M-25-22 (OMB, April 2025) directs agencies to require ongoing testing with validation data that “should not be accessible to the vendor,” and vendor tests “detailed enough to be independently verified or reproduced, if practicable” — neither alternative method is preferred over the other; one of them is required.
Reporting differs too: the 2025 federal inventory has two tiers — a consolidated COTS (software bought as-is, not built) certification for the suite-AI class, and individually reported systems — and 46 agencies submitted consolidated COTS data (OMB). High-impact is never consolidated, so a system determined high-impact — Meridian’s “consequential” is the house tier, and the federal determination is a separate, documented finding (1.3’s written determination) — would be reported on its own, with its testing, independent review, and appeal path named.
The handoff artifact is the protocol with its observations, the vendor-run declaration, and the register’s reporting tier. What is not equivalent: a commercial query-and-observe run is a practice the team chose and can drop; the federal one is a minimum practice, and consolidated reporting exists only for the bought-as-is tier — a system determined high-impact cannot be certified in a batch. Where Meridian’s task order incorporates those terms, the route used is on the record in the memo’s own words before anyone asks.
Stop and escalate when the vendor will neither expose an interface for the set nor run the sealed data. Then no route exists, the system has no evaluation, and Lesson 1.3’s exits apply: it does not deploy — or, if it is already running, it enters 1.3’s safe-discontinuation exit, where continued federal use needs a pilot exemption or a waiver and commercial use needs a remediation deadline the risk acceptor signs — not a benchmark table nobody at Meridian can reproduce.
The ATS vendor ran Meridian's sealed applicant set through its ranking model and returned a results table with a one-page description of how it ran. What may the record say?
Practice status — among organizations whose register is mostly provider-controlled systems, commercial and federal
| Practice | Status | Also called |
|---|---|---|
| query-and-observe on a sealed set | required where M-25-21 flows down; common baseline elsewhere | black-box acceptance testing |
| vendor-run results declared as vendor-run | strong optional | third-party test attestation, labeled |
| reproducibility stated on the record | required where M-25-22 flows down; emerging elsewhere | reproducible vendor test |
| collection route (query-and-observe / vendor-run) recorded beside the evidence level | strong optional | provenance tag on the evidence row |
| consolidated reporting for bought-as-is tools | federal-only mechanism | consolidated COTS certification |
| consequential systems reported individually | required in federal reporting; reference-shop elsewhere | high-impact never consolidated |
Scale: required | common baseline | strong optional | reference-shop (seen only at organizations that publish their own practice) | emerging
Key takeaway
For the systems you cannot open, two routes exist and both are recorded as what they are: query-and-observe on the sealed set through the product’s own interface, with what it cannot claim written down, and sealed data handed to the vendor with the result declared vendor-run, method attached, reproducibility stated. Lower assurance is recorded as lower assurance. The federal memos name these routes, direct agencies to require validation data the vendor cannot reach, and never let a consequential system be certified in a batch; where no route exists, the system has no evaluation and the exits from 1.3 apply. Module 4 turns to attacking the evidence you now have.
LEADERSHIP DECISION accept a lower-assurance route for a system you cannot open only when the record says so in those words - and refuse it for a consequential system with no route at allPRACTITIONER ACTION run the protocol by hand on the sealed set; declare every vendor-run result with its method and its reproducibilitySUCCESS MEASURE zero provider-controlled systems on the register carrying a vendor table as their only evidence - an audit finding avoided