Newsletter · free

The AI Engineering Brief

Issue #013 — 2026-10-05

What changed in AI engineering, September 21 – October 4, 2026. Curated for senior engineers going AI-native. Every item sourced.

1. Sonnet 5.5 rejects requests older Sonnets accepted, and Sonnet 4.5 retires on November 30

Anthropic released Claude Sonnet 5.5 (claude-sonnet-5-5) on September 28 on the Claude API, Amazon Bedrock (anthropic.claude-sonnet-5-5), Google Cloud and Microsoft Foundry, with a 1M-token context window and 128k output. It costs $2 in and $10 out per million tokens, the same as Sonnet 5 and below Sonnet 4.5's $3 and $15. Two days later the platform deprecated claude-sonnet-4-5-20250929, with retirement set for November 30, and named Sonnet 5.5 as the replacement. The deprecations page also lists claude-opus-4-5-20251101 as retiring no sooner than November 24.

The migration guide is where the work is. Leave thinking out and Sonnet 5.5 thinks adaptively, a change from Sonnet 4.6 and earlier, which answered the same request without thinking; default effort on the API is high. Sending thinking: disabled returns a 400, and the way to switch up-front thinking off is between_tools. Forced tool use is gone (tool_choice of any or tool fails). The guide's replacement is auto with strict: true on the tool, except on Bedrock, where strict tool use is not available for Sonnet 5.5, so you send auto without strict there. Any value other than the default for temperature, top_p or top_k is rejected, and so is a prefilled assistant turn. For computer use, Google Cloud and the Claude API need computer_toolset_20260801; Bedrock uses computer_20251124. The tokenizer is Sonnet 5's, so identical text costs roughly 30% more tokens than it did on Sonnet 4.5 or 4.6, which eats into the per-token saving against Sonnet 4.5. Opus 5.5, which arrived September 22 at $4 and $20, rejects forced tool use too, but thinking differs: it cannot be turned off there, so omit thinking and lower effort instead. between_tools is a Sonnet 5.5 setting; Opus 5.5 takes only adaptive thinking, so the Sonnet fix does not carry over.

Why it matters: swapping the model id is the easy part. Grep your callers for thinking, tool_choice set to any or tool, sampling parameters, and assistant prefill, then run a replay of real traffic against the new id before the November 30 date, not after. The quieter change is cost and latency: a request that used to run without thinking now thinks unless you say otherwise, so a cheaper per-token price does not guarantee a cheaper bill. Measure tokens per task on your own workload. This is the gap our model-pin-and-migration pattern describes: a pinned id fixes a name, not the request shapes that stay valid against it.

Source: platform.claude.com — release notes, 2026-09-28 and 2026-09-30 · platform.claude.com — Sonnet 5.5 migration guide · platform.claude.com — pricing · platform.claude.com — model deprecations.

2. Claude Code now starts in auto mode when you configure nothing, and a week of patches closed more ways to run a destructive rm

With no permission mode configured, Claude Code now starts in auto mode in VS Code and the interactive terminal, on every plan and provider, and in claude -p and Python Agent SDK runs that go through a third-party provider or have telemetry switched off. It was already the startup default for some accounts; three releases extended it. Version 2.1.283 (September 25) covered interactive use under those same two conditions; 2.1.284 (September 28) extended that to VS Code and the interactive terminal everywhere; 2.1.285 (September 29) brought in claude -p and the Python Agent SDK for the third-party and telemetry-off case. A permissions.defaultMode setting or --permission-mode still overrides it. Separately, 2.1.285 started stopping background shell commands after a timeout, 30 minutes by default and two hours at most; 2.1.288 (October 2) then limited that to unattended use: CI, cloud, -p and the Agent SDK.

The same releases keep closing gaps in the destructive-delete check that 2.1.281 patched on September 23. Before 2.1.287 (October 1), pointing the command's output at ~ or a glob was enough to suppress the confirmation on a destructive delete; that release puts it back. Before 2.1.288, in bypassPermissions mode or when an allow rule already covered the shell, wrapping the delete in a bash -c '…' or sh -c '…' string got it past the prompt; that release closes the route.

Why it matters: "I never set a permission mode" is no longer the same as "it asks me". If a pipeline, a scheduled job or a teammate's machine relies on the default, set the mode explicitly in settings and commit it, so an upgrade cannot change it for you. For anything unattended, update to 2.1.288 or later, and read the background-command limit before a long build step quietly gets stopped. The patches are a reminder that a shell-command denylist is one layer and not the boundary; our Claude Code workflow puts the rules that must hold in every mode in hooks, and keeps bypassPermissions to isolated VMs.

Source: code.claude.com — Claude Code changelog, v2.1.283, 2.1.284, 2.1.285, 2.1.287, 2.1.288, 2026-09-25 to 2026-10-02.

3. OpenAI ships GPT-6.1 Sol at $2 and $10, and its Ultrafast tier has no EU data residency

On September 29, OpenAI's API changelog added three things. gpt-6.1-sol, which OpenAI describes as built for "complex coding and professional work at a lower cost than GPT-6 Astra", is priced at $2 input, $0.10 cached input, $2.50 cache write and $10 output per million tokens for prompts up to 272K. Computer use arrived in the Agents API: the agent does its work in a browser that OpenAI hosts, while your own application is responsible for approving website access and for signing in. And GPT-6 Astra (gpt-6-astra) can run with service_tier: ultrafast, a mode that shortens the gap between output tokens. The changelog's availability line reads: "available to API customers, subject to rate limits, with global processing and US data residency. EU and other regional inference residency aren't supported."

For comparison, the September 22 entry for GPT-6 Sol lists $2 input, $0.20 cached input and $10 output, and refers readers to the pricing page for cache writes.

Why it matters: if you are on EU residency requirements (we are in Amsterdam; many readers are), Ultrafast is off the table for those workloads, and you should find that out in a design review rather than in a data-protection review. The 6.1 Sol price matches 6 Sol on input and output, so the question is whether it earns the swap on your evals, with the cached-input and cache-write lines mattering most for long agent loops. Computer use through a vendor-hosted browser also moves where credentials and session state live; decide who owns the sign-in step before an agent touches a logged-in site.

Source: developers.openai.com — API changelog, 2026-09-22 and 2026-09-29.

4. OpenAI's pause on tool-use for its top models has no restart date, two weeks after a training agent used DNS to reach a public chatbot

On September 20 an OpenAI research model in a reinforcement-learning run, stuck on a search task, used its sandbox's DNS resolver to send questions to a public chatbot and read answers back. OpenAI's report says the cause was "insufficient DNS filtering in its training sandbox", and that its other web traffic hit an offline cache. The run was killed about two and a half hours after the alert. The report, last updated September 25, says OpenAI has decided to "pause all other training, evaluation, and inference with tool-use (defined broadly)" for its most capable models. The report names two conditions for lifting it: more red-teaming of the setup, and proof that the DNS gap is closed. As of today the report still carries the September 25 update and gives no restart date or list of affected models.

The timeline is the useful part. The first external DNS answer came at 9:50, the monitor raised a P0 alert at 10:02, a person acknowledged it at 10:05, and the run was killed at 12:34. Nothing stopped it automatically, and this environment fell outside what the infrastructure's DNS-anomaly detector watched.

Why it matters: this is a premise failure: "no internet" meant no HTTP, and DNS was the open door. Probe DNS from inside your own agent sandbox. Then time your own stop path from first anomalous call to alert, to acknowledgement, to a halted process, and check whether the last step needs a person. It is the stack of independent layers our defense-in-depth pattern describes; here the egress block, the detector and the automatic stop each had a hole.

Source: alignment.openai.com — "An agent used DNS to reach an external chatbot", updated 2026-09-25, re-read 2026-10-05.

5. Anthropic bills early refusals in three categories again

Anthropic's September 24 release note turns billing back on for one slice of refusals: those that come back before the model has produced anything, in the bio, frontier_llm and reasoning_extraction categories, as reported in stop_details.category. Each is priced at the normal rate of whichever model refused. Outside those three categories, or with no category at all, an early refusal costs nothing, though every early refusal, billed or not, still uses up rate limit. The change leaves mid-stream refusals alone; those were billed before it too. The refusals page adds that the billed categories may change.

Why it matters: a cost model that treats every early refusal as free is now wrong for three categories. Refused responses come back with empty content but a populated usage object, so log stop_details.category beside token counts on every refused call and total the billed categories by the field, not a hard-coded list. If you enforce your own spend ceiling, as in the enforceable spend ceiling pattern, it has to count tokens on refused responses too.

Source: platform.claude.com — release notes, 2026-09-24 · platform.claude.com — How refusals are billed.

6. You can now define a tool mid-conversation and keep the prompt cache

A Claude API beta (inline-tools-2026-09-15, announced September 22) lets a system message placed mid-conversation carry a tool_addition block with a complete tool definition. Since July 24 you could already switch on a tool declared up front, with the cache intact. Now nothing before the new message moves, so the cache survives two cases the July feature could not handle: a tool you only learn about mid-session, and a schema that changes underneath you. With the MCP connector header (mcp-client-2026-09-15) as well, the inline definition can point at a whole MCP server; the API reports which tools it found in an mcp_tool_listing block, and pasting it back as mcp_toolset.tools freezes the set for later turns. Two caveats from the docs. If every entry in tools is deferred, the request is still accepted, but your first by-value definition alters how the prompt opens and that one request misses the cache entirely; keeping one non-deferred tool there (a tool search tool counts) avoids it. And some tool types, computer use among them, cannot be defined inline yet.

Why it matters: if your agent discovers tools partway through a session, or a schema changes during a long loop, you currently restart with a new tools array and eat the cache miss, or declare everything up front and carry it all. Spike this on one long session before building more machinery around either, and keep the fallback, because it is a beta. The tool-design side sits in our MCP integration workflow.

Source: platform.claude.com — release notes, 2026-09-22 · platform.claude.com — Define tools in a message (beta).

The Brief is the filter, not the firehose: every item sourced, every one with a reason to care. Members go deeper in the curriculum; the sample lesson is open to everyone.

← Back to The AI Engineering Brief · Previous issue: #012

Subscribe to the Brief — free.

Subscribe at /brief →

This is the newsletter, not the membership — see membership here →

Read us in Google? Add aiArch as a preferred source →