Pattern · last reviewed 2026-10-03

Pinning a model is not pinning its behaviour

What this pattern is

A pinned model id fixes a name in a config file. It does not fix the serving stack behind that name, the defaults applied to it, or the tokenizer that decides what your prompt costs — so the migration you pinned against can happen without the id ever changing. What actually needs pinning is a measurement: an eval that tells you the behaviour moved. Build the id as a stable role→id mapping (see the model router), indirect the wire-level provider slug behind it, and put a scored eval suite in the release path so a change in what the id does is caught the same way a change in what your code does would be.

Context & problem

A team pins a model id — claude-opus-4-8, say — because "pinned" reads as "stable." It is stable in exactly one respect: Anthropic's own deprecation policy says an active model id will not be retired without at least 60 days' notice, delivered by email and in the documentation. That is a real guarantee and it is worth having. It is also not the guarantee most teams think they bought.

The id names a model. It does not freeze the API surface around that id, the infrastructure serving it, or the numbers a tokenizer assigns to your prompt — and all three can move while the id you pinned sits unchanged in your config. Anthropic's own parameter table makes this concrete: as of the model-deprecations page fetched for this pattern, temperature, top_p, and top_k are deprecated on Claude Opus 4.7 and later and now return a 400 error when set to a non-default value — on ids that were already active, that nobody retired, that a pinned config still names correctly. A team that pins the id and calls the job done has pinned against the wrong failure mode.

Forces

  • What retirement covers vs what it doesn't. Anthropic's lifecycle terms — Active, Legacy, Deprecated, Retired — and its 60-day notice period are about the id disappearing. They say nothing about the id's behaviour while it remains Active, because that was never the promise.
  • A default isn't visible in the id string. Anthropic's own docs say effort defaults to high on Sonnet 5 and Opus 5 — on every route, first-party API included — but a pinned id doesn't print that, and it's easy to assume a default you first noticed on one route is specific to that route. Our own guard was scoped that way at first and had to be widened; detailed below.
  • Stability vs staying current. An indirection layer plus an eval suite is a standing cost that buys you a warning, not a guarantee; skip it and you find out from a user complaint instead of a red test.
  • What you can verify vs what you can only hope. Some mitigations are provable with a live call; others — like whether a wire-level flag actually does what its name says — stay unverified until you spend to check, and the honest version of this pattern says so rather than claiming a fix that was never confirmed.

The pattern

Treat "what model does this call use" as two separate questions with two separate answers. The first is identity: a stable role→id map, one function, so a model swap is a one-line edit (this is the model router pattern). The second is behaviour: does the id still do, today, what your product depends on it doing. The router answers the first question. Only an eval answers the second, and an eval only answers it if it actually runs before a ship — as the release gate the eval-harness gate pattern already describes, pointed here at model drift rather than at your own prompt changes.

Between the two, put an indirection boundary: the canonical id your code reasons about should not be the literal string sent on the wire. A provider slug, a gateway route, or a partner-platform mapping can all drift independently of the model itself, and when one of them does, you want that to be a one-file diff, not a grep across every call site.

What a pinned model id actually covers A pinned id protects against retirement, which carries a notice period. It does not protect against parameter deprecations, a default the id ships with, or tokenizer drift on the same id — only an eval harness catches those. pinned model id unchanged in config API parameters e.g. temperature 400s the id's own defaults e.g. thinking, id-wide tokenizer cost per prompt eval harness the actual pin
The id stays fixed on the left. Three things can still move underneath it — parameters, the id's own defaults, the tokenizer. The eval harness on the right is the only box that actually notices.

Reference implementation notes

Anthropic

Anthropic's own model-lifecycle documentation is explicit about what a pinned id buys you: four states — Active, Legacy, Deprecated, Retired — and "Anthropic notifies customers with active deployments for models with upcoming retirements, providing at least 60 days' notice before model retirement for publicly released models." That is a real, citable guarantee against the id disappearing out from under you.

It is a narrower guarantee than it sounds, because the same page documents behaviour changing on an id that stays Active. Its parameter-deprecation table states: "temperature, top_p, top_k — Deprecated (Claude Opus 4.7 and later) — Returns a 400 error when set to a non-default value on Claude 4.7 and later models" — with the recommended fix being to drop the parameter and steer behaviour through prompting instead. Separately, the same page notes the Python SDK (v1.0 and later) removes those three parameters from the request types entirely, so old code that sets them raises a client-side TypeError before the request is even sent. Neither event retires a model id. Both change what code written against that id needs to do.

AWS (Bedrock)

Amazon Bedrock runs its own, separately-dated lifecycle on top of Anthropic's: a model on Bedrock is Active, Legacy, or End-of-Life, and — this is the guarantee closest to what teams assume "pinned" means — "once a model launches on Amazon Bedrock, it will remain on Amazon Bedrock for at least 12 months before the EOL date," with a further minimum of 6 months in the Legacy state before EOL, per Bedrock's model-lifecycle documentation. The same page is explicit that this is a Bedrock-specific clock: "Model lifecycle dates on this page are specific to Amazon Bedrock and may differ from dates published by model providers (such as Anthropic or Cohere). For Amazon Bedrock usage, only the dates on this page apply." A team running the same nominal model through both Anthropic's API and Bedrock is tracking two independent retirement calendars under one id, and the Bedrock guarantee — like Anthropic's — covers the id's continued existence, not what it does while it exists.

Trade-offs

  • The eval suite is the actual cost. An indirection layer is cheap; a scored eval suite that would actually catch a behaviour drift is not — it needs real cases, a scoring function, and someone reading the score when it moves. Skipping the eval and keeping only the indirection buys you the wrong half of the pattern.
  • A stale eval suite is worse than none — it launders drift as "unchanged." If the suite encodes last quarter's answers as ground truth, a model that got worse in a way the suite doesn't test still shows green.
  • Two independent lifecycle clocks if you run on more than one platform. Anthropic's own retirement schedule and a partner platform's (Bedrock's, Vertex's) can diverge for the same nominal model — verify against the platform you actually call, not the vendor's headline page.
  • A wire-level mitigation you have not tested is not a mitigation. Sending a flag that is supposed to suppress a behaviour is not the same claim as having confirmed it suppresses the behaviour — see the as-built section below for a case where we shipped the first without the second, on purpose, and said so.

When not to use this

Skip it early, on a product still finding its shape. Pinning plus an indirection layer plus a shadow-eval suite is a standing cost, and it buys a stability you may not want yet: pin at the wrong moment and you sit out a generation of cheaper, better models while your eval suite encodes last quarter's answers as ground truth. If you ship weekly and would take a free quality upgrade on sight the moment it's available, float the version and spend the effort on the harness instead — the harness is what makes the pin optional later, and it is the piece worth building first regardless of which way you land on pinning itself.

As-built evidence

aiArch pins by role, not by literal wire string. modelFor() in src/lib/llm.ts returns a stable canonical id per role (hint, tutor, eval, grader, safeguard); a separate wireModel() maps that canonical id to the OpenRouter slug actually sent on the wire, through a small lookup table the code comments explicitly mark "re-confirm on release." The two-function split exists because the mapping has already drifted once in production: Sonnet 5 replaced Sonnet 4.6 as the tutor role's id on an unchanged request shape, and the file keeps the retired slug mapped for historical cost accounting rather than deleting it outright. It happened again on 2026-10-03, when tutor and grader moved to Sonnet 5.5 and eval to Opus 5.5: a role-id edit plus new pricing and wire-slug entries, with Sonnet 5 and Opus 5 left mapped as legacy.

The behaviour-drift half of this pattern is not hypothetical here either. Our coach began returning empty replies that still billed, with no id change and no code change: claude-sonnet-5 ships with adaptive thinking on by default, and reasoning tokens count against max_tokens — a hard-thinking turn burned the whole output budget with zero visible content, a successful call that said nothing. We first hit this via OpenRouter serving the id through Amazon Bedrock (2026-07-08), and the fix we shipped was scoped to that route, reading at the time as a Bedrock-specific workaround. It wasn't: Anthropic's own docs state effort defaults to high on the first-party Claude API too, for both Sonnet 5 and Opus 5, so the guard was widened (2026-07-26) to run on every Anthropic route the code sends, not only the one where it was first caught — a narrower fix would have quietly regressed the moment traffic moved off Bedrock. We send reasoning: { effort: 'none' } on the Anthropic slugs that accept it, such as Haiku 4.5 (buildChatBody, src/lib/llm.ts). The two 5.5 slugs we route tutor, grader and eval to reject that value, so since 2026-10-03 they get 'low' instead — thinking still runs there, bounded by effort, and still counts against max_tokens. The honest status of that fix is recorded, not assumed: test/llm.test.ts asserts the wire shape under a title ending "(wire shape, not efficacy)" — deliberately not claiming the parameter works, because whether it actually suppresses thinking on Haiku 4.5 is still an open, unsettled question in our own tracker (Q-353). Settling it takes one live gateway call reading the thinking-token count back off the usage object, which is real spend, and we have not spent it for Haiku. On 2026-10-03 a live eval run did read reasoning_tokens on the Sonnet 5.5 slug (peak 178 on a tutor turn), which measures the 'low' setting there but says nothing about 'none'. What we shipped instead as the control that actually holds is runCoachLoop yielding an in-band "ran out of reply budget" message on any turn that produced no visible text (src/lib/coach.ts) — visible and survivable whether or not the wire parameter does anything at all. That ordering — verify what you can, and make the thing you can verify the one holding the weight — is the transferable part.

The eval harness half of the pattern runs as npm run eval (test/coach.eval.test.ts, test/grader.eval.test.ts, test/practical.eval.test.ts), scored suites gating the same release path a prompt change would go through — the mechanism the eval-harness gate pattern names generally, pointed here at "did the model change under me" rather than only "did my prompt change break it." It is a role-scoped harness for our own call sites, not a general Anthropic-fleet drift monitor — the 2026-08-24 currency-watch sweep that found the temperature/top_p/top_k deprecation live on our own tracked ids (Opus and Sonnet tiers, unretired, unchanged) is what a currency-watch process catches; the eval suite is what would catch a regression in what those ids actually produce.

As-built: the Sonnet 5.5 move failed our behaviour gate

Moving the tutor role to Sonnet 5.5 on 2026-10-03 was a role-id edit, as the section above says. The id change was the easy part. The live coach gate (npm run eval:live, test/coach.eval.live.test.ts) is what told us the behaviour had moved.

What the gate is, and what it is not

The gate runs 35 scripted coach cases against the real model. Each case is scored by deterministic checks in src/evals/coach.eval.ts: a stop-reason check, checks on which tools were called, string and pattern checks on the reply, and one rule that the reply must end on a re-check question or a report-back invitation (closesWithRecheck). It runs up to three times, passes when two runs reach 80% and fails when two fall short. That median-of-3 rule replaced a single-run gate that had been a coin flip at the boundary: identical code scored 74.1% on one run and 81.5% on another. In this suite the input-injection judge is stubbed out (judgeFn returns unavailable), so the red-team cases reach the coach's own defences. The judge has a separate live eval. So 77% does not mean the tutor teaches 23% worse. It means that in 23% of cases the reply broke a rule we had written down, and the rule is a proxy for good teaching, not a measurement of it.

AttemptRun scores (35 cases each)Verdict
Sonnet 5.5, unchanged prompt74.3%, 82.9%, 74.3% (77.1% pooled)Fail: two runs under 80%. Sonnet 5 had passed the same gate.
Sonnet 5.5, after one prompt round and a larger reply budget91.4%, 94.3%Pass: two runs over 80%, the third not needed.

Dump the replies before you touch the prompt

A score tells you that behaviour moved, not why. Two explanations predicted opposite fixes. If thinking was eating the output budget and cutting replies off, the fix was more tokens and a prompt edit would do nothing. If the replies were complete but shaped differently, the fix was the prompt and more tokens would do nothing. Running with EVAL_DUMP=1 (test/evalDump.ts) writes one line per case: the full reply as scored, the finish_reason of each model call, and the raw usage including reasoning tokens. Reading the three runs separated the two.

  • Mostly style. Apart from one truncated reply and one case the output filter blocked in all three runs, every failing case ended with stop and carried a full reply. Reasoning ran on 37 of 168 model calls and peaked at 178 tokens. Of 15 closesWithRecheck failures, 13 had a sentence after the last question mark ("Answer in your own words.", "Once you answer, I'll…"), 1 phrased the re-check as an instruction with no question mark, and 1 was the truncated reply below. Scope pivots said "the course" or "your weakest objective", where the checks look for words such as "module", "lesson" or "curriculum".
  • One real truncation. One case in one run stopped on length at exactly 1024 completion tokens: 117 reasoning tokens plus about 907 visible. Its two sibling runs finished at 889 and 928 tokens. The largest untruncated reply was 928 tokens and peak reasoning was 178, so 1024 had no room for both.

The fix matched the evidence. One prompt round added a hard rule that a question the coach asks is the last sentence of the reply and ends in a question mark, required scope pivots to name the module, and made the canned reply for a blocked output point back to the lesson. The tutor's maxTokens went from 1024 to 2048. We changed no check and edited no failing case's input, which is the harness's standing policy: when the gate fails, change the thing under test.

The wire change underneath

OpenRouter lists reasoning as mandatory on both 5.5 slugs and rejects effort: 'none' on them, so buildChatBody sends 'low' there (src/lib/llm.ts). Thinking still runs and still counts against max_tokens, which is why the larger budget mattered.

What to take from it

  • A model upgrade is a behaviour change. Re-run the behaviour gate on every model move, including a move to a newer version of the same family at the same price. A passing typecheck and unit suite told us nothing here. The tutor did not ship until the gate passed.
  • Dump replies before you tune. Decide between truncation and style from the artefact, not from a hunch.
  • Keep the bar frozen. Change the prompt, the budget or the model, not the checks. A gate you loosen to pass is no longer a gate.

Limits of this evidence: 91.4% and 94.3% are two runs after one round of tuning against that same dump, so the cases we tuned on are the cases now passing. Nothing here is a held-out set. The checks are partly string-shaped, so some of the lift is the reply conforming to a format we can measure. The live gate is also real spend, so it runs by hand on a model move rather than on every commit. For the general mechanism, see the eval-harness gate.

Changelog

  • 2026-10-03 — Added the as-built section on the Sonnet 5.5 move: the live coach gate scored 77.1% before one prompt round and a larger reply budget, 91.4% and 94.3% after.
  • 2026-08-25 — Initial publication. Anthropic model-deprecation policy and parameter-deprecation table verified against the live docs; AWS Bedrock model-lifecycle policy verified against the live user guide.
Sources & provenance
  • Anthropic model lifecycle terms (Active/Legacy/Deprecated/Retired), the 60-day retirement notice commitment, and the temperature/top_p/top_k parameter-deprecation table (400 on non-default values, Claude Opus 4.7 and later; removed from the Python SDK v1.0+ request types): Anthropic model deprecations docs, fetched 2026-08-25.
  • AWS Bedrock model lifecycle (Active/Legacy/EOL states, minimum 12 months on Bedrock before EOL, minimum 6 months in Legacy before EOL, Bedrock-specific dates independent of the model provider's own schedule): AWS Bedrock model lifecycle docs, fetched 2026-08-25.
  • The Bedrock-served-Sonnet-5/extended-thinking incident, the shipped reasoning: { effort: 'none' } mitigation, and its unresolved-by-decision efficacy status: this platform's own src/lib/llm.ts and src/lib/coach.ts, and Q-353 in our internal engineering tracker, stated as first-person practice.
  • The Sonnet 5.5 gate failure and fix (run scores, reply dump findings, prompt round, maxTokens 1024 to 2048): this platform's own live coach gate, test/coach.eval.live.test.ts, src/evals/coach.eval.ts, test/evalDump.ts and the as-built tutor-gate diagnosis of 2026-10-03, stated as first-person practice. The mandatory-reasoning wire rule: THINKING_MANDATORY in src/lib/llm.ts, read against OpenRouter's model list 2026-10-03.
  • The role→id split and the wire-slug indirection: modelFor()/wireModel(), src/lib/llm.ts. The eval harness: test/coach.eval.test.ts, test/grader.eval.test.ts, test/practical.eval.test.ts, run via npm run eval.

Deprecation dates, retirement notice periods, and provider lifecycle terms are the facts here most likely to drift — re-verify against the live pages before relying on any figure newer than 2026-08-25. Corrections: hello@aiarch.dev.

Learn the cost and reliability engineering underneath the pattern.

aiArch teaches model selection, migration, and production agent architecture by building — on a platform that pins by role, routes the wire slug through one file, and records when its own mitigation is unproven instead of overstating it.

Free with an account: Track 0 (The Mindset Shift) and the Track 0 coach allowance, no card · every claim cited · full curriculum with membership