The Version Boundary
Last reviewed · content updated
IntermediateWhat you'll learn
~20 min- Tell an alias from an immutable version identifier, and record which one each of the ten systems actually has
- Write the version-boundary line every release-decision record carries, including the announced retirement date
- Explain, from the published drift evidence, why a re-run is not optional when the boundary moves
Every provider name, identifier, and date below is true as of August 2026 and re-verified monthly — evidence with sources, not a rule about any vendor.
Before the detail — Decision: pin the version the evidence describes, or record that you could not. Outcome: a retirement notice becomes a scheduled trigger, not a surprise. Artifact: the version-boundary line in the record. Status of what follows: reusable guidance; dated as of August 2026.
Prompt first: draft ten boundary lines, pin none
Here is Meridian's AI inventory [paste scripts/t7-substrate/data/ai-inventory.yaml].
For each of the ten systems, draft the VERSION BOUNDARY line fromthe model_provider field alone: provider vendor, in-house, or appliance immutable id the exact version string, or the weight digest for local weights, IF stated - else NEEDS-OWNER alias if the field names only a family, a product, or a "latest"-style name: ALIAS - NOT PINNED, plus one line on what that means for the evidence license NEEDS-OWNER notice date NEEDS-OWNER retirement date NEEDS-OWNER unless announced in the field
Do NOT supply a version, digest, license, or date from what youknow about any vendor.Ten lines, mostly NEEDS-OWNER — which is the lesson. Not one model_provider field states an immutable identifier: “model family not disclosed in contract,” “commercial code-assistant subscription,” “productivity-suite built-in AI.” The agent makes that visible; the owners go to the console, the contract, or the weights file and replace each NEEDS-OWNER with a value or with “cannot be pinned, and why.”
A name is not a version
Lesson 1.2’s register has a version-boundary object; a name off an invoice cannot fill it.
Aliases hot-swap. Google’s documentation states its -latest names are hot-swapped with every new release: your configuration string stays, the model behind it changes. Default snapshots are not the newest. OpenAI’s documentation states that its gpt-4o name resolves to a default snapshot that is not the most recent one published. Deployments auto-upgrade unless pinned. Azure’s documentation states that Standard deployments auto-upgrade. Fixed weights, moving behavior. Anthropic’s documentation states its current identifiers are weight-pinned while routing, safety classifiers, and sampling can still change — a fixed weight ID (the exact set of trained parameters) is the strongest thing a hosted provider offers, and even that is not the whole system that answers.
For local weights — MU-AI-004’s model, MU-AI-009’s classifier — the immutable identifier is the file’s digest, the hash Lesson 2.3 puts into the bill of materials. That one you can pin.
Retirement clocks, as of August 2026
Hosted models retire, and the clocks are short. Every row comes from the provider’s published deprecation, retirement, or policy page.
RETIREMENT CLOCKS as of August 2026; provider-published; re-verified monthly legend: model - what happened - source
Mistral Small 3.1 notice Nov 6 2025, shutdown Nov 30 2025: 24 days provider deprecation page gemini-2.5-flash-preview retired after 5 months provider deprecation page Azure gpt-5-chat retired at 10.7 months provider retirement page claude-opus-4-1 retired at 12 months provider deprecation page gpt-5-2025-08-07 retiring at 16 months provider deprecation page
policies: OpenAI GA at least 6 months notice, previews "as short as 2 weeks"; Anthropic at least 60 days; Azure Standard deployments auto-upgrade provider policy pagesThe record names the retirement date and Lesson 6.1 turns it into a trigger; this lesson states no “flag if retirement is within so many months” rule, because no source supports one. A preview model is not a production boundary: a two-week notice cannot be absorbed by a re-evaluation cycle longer than two weeks.
Why the re-run is not optional
The same name, months apart, is not the same behavior. Chen, Zaharia, and Zou’s 2023 study measured GPT-4’s accuracy at identifying prime numbers at 84% in March 2023 and 51% in June 2023, under the same model name. Whatever one thinks of the task, the shape is the point: March’s evidence said nothing about June.
Grounded Answers 6.1 (a separate training in this series) taught it — the answer names its versions. Here the Module 3 evaluation and Module 4 attack evidence attach to the boundary line; when the boundary moves — alias swap, auto-upgrade, retirement, a routing change under a fixed ID — that evidence describes a system that no longer exists.
What a good card looks like
Lesson 2.1 said the attempt curve is the bar and the cards that publish one set it. Two, as published claims — the provider’s numbers, never Meridian’s evidence. Anthropic’s Claude Opus 4.5 system card (November 2025) reports the Gray Swan Agent Red Teaming benchmark: indirect prompt injection (instructions hidden in content the model reads) succeeded against Opus 4.5 Thinking 0.3% of the time at a single attempt, 1.8% at ten, and 25.0% at a hundred — the curve itself is the point, and the primary card is what the packet cites. Read the chart, not the card: the same document carries a second chart on the facing page that folds direct injection and jailbreaking into the same benchmark, where the same model reads 4.7%, 33.6% and 63.0% — quote that one as the indirect rate and you have overstated a single attempt by more than fifteenfold. OpenAI’s GPT-5.6 system card publishes an indirect figure of 3.77%, measured against its own automated red-teamer with the attacker holding a single message, and no attempt scale at all.
The first is what a curve looks like: the figure at a hundred describes a persistent attacker; the figure at one describes exposure to a single attempt, which is a real number answering a much smaller question. The second cannot be set beside it — different benchmark, different attacker, no k — so a packet that prints them as a two-row comparison has invented a gap neither vendor measured. Neither is Meridian’s number — Module 4 produces that, at Meridian’s attempt budgets (how many tries the test allows before it stops and scores the run).
The risk: a single-attempt figure read as the curve understates contract risk by an amount nobody measured.
The boundary line
Eight things: provider; immutable ID, or the digest for local weights; the alias you could not pin and what it means; the endpoint and deployed configuration as the provider exposes them (routing, safety classifiers, revision); the date observed; license; notice date; retirement date if announced. MU-AI-004’s line pins a digest; MU-AI-005, a vendor-hosted LLM behind an enterprise API, cannot:
VERSION BOUNDARY MU-AI-005 Procurement Document Summarizer provider vendor-hosted general LLM, enterprise API immutable id NEEDS-OWNER - contract names a product, not a snapshot alias ALIAS - NOT PINNED: evidence attaches to whatever the product name resolves to on the day of the run endpoint NEEDS-OWNER - routing, safety classifiers and revision as the provider exposes them date observed NEEDS-OWNER - the day the boundary above was read license commercial terms, dated (packet, Lesson 2.1) notice date NEEDS-OWNER - the change-notice clause from 2.1 retirement date NEEDS-OWNERA line that says NEEDS-OWNER is honest; one that says a product name is a finding.
Stop and escalate when a production system’s boundary resolves only to an alias and the vendor cannot supply an immutable identifier or a written notice commitment — the line records “cannot be pinned,” and whether Meridian accepts evidence attached to a moving target is the risk acceptor’s decision in writing, before Module 3 spends hours on it.
MU-AI-001's portal assistant runs against a provider alias, and the alias moved last week. The evaluation passed in June. What is the status of that June evidence?
Key takeaway
A name alone does not prove an immutable version. Aliases hot-swap, snapshots lag, deployments auto-upgrade, and a fixed weight ID can still answer differently; hosted models retire on clocks measured in weeks to months. The drift study is why the re-run is not optional; the published attempt curves are the provider’s numbers, recorded as claims. Every record carries the boundary line, and 6.1 turns its retirement date into a trigger. Lesson 2.3 turns to the lineage you can actually verify: what is new about weights.
LEADERSHIP DECISION accept that hosted models move on the provider's clock; fund the re-evaluation each move forcesPRACTITIONER ACTION write the boundary line for every system; pin an immutable ID or digest where one exists, record ALIAS - NOT PINNED where none doesSUCCESS MEASURE every approved record names a boundary; zero evaluation results attached to a name that moved since the run