Sonnet 5.5 keeps its reasoning for one model and one account
Switch to a model or an API account that can't read the earlier thinking, and the request still succeeds. The reasoning those turns built up is dropped before the model reads it, with no error and nothing on the bill. You only find out if you asked the API to report it.
Claude Sonnet 5.5 shipped on September 28 at the same prices as Sonnet 5, and Claude Code v2.1.284 made it the default Sonnet on the Anthropic API the same day. Anthropic lists five ways code written for Sonnet 5 breaks on it. Forced tool_choice returns 400 and computer use moves to the toolset on the Claude API and Google Cloud, both familiar from Opus 5.5. The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors. Turning thinking off works differently from Opus 5.5: disabled returns 400 and the replacement is between_tools, at high effort or below. And thinking blocks are now tied to the model and the account.
I wrote last week that a 400 is the good kind of break. The fifth change is the one that doesn't raise one, and it quietly moves the unit a model router should work in: from the turn to the conversation.
Which models can read Sonnet 5.5's thinking?
None besides Sonnet 5.5 itself. The docs are explicit: Sonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models, but not from Opus 5, Opus 5.5 or any Fable or Mythos model, and "no other model reads thinking blocks from Claude Sonnet 5.5."
Look at which pair that cuts. The routing this lineup invites is Sonnet 5.5 for routine turns and Opus 5.5 for the hard ones. Those two share nothing in either direction. Opus 5.5 reads Opus 5 and earlier Opus, Sonnet and Haiku blocks, so Sonnet 5 to Opus 5.5 keeps its reasoning. Sonnet 5.5 to Opus 5.5 does not, and neither does a route back down to Haiku 4.5.
What happens to a block the model can't read?
The API removes it before the prompt reaches the model, and the request succeeds. Dropped blocks aren't billed and don't count toward input_tokens. The text and tool_use blocks stay, so the model still sees what was said and done. It just answers that turn without the reasoning behind it.
The loss is per request, not permanent. The API never edits your messages array, so the next request that goes back to a model able to read those blocks gets them back. Anthropic's advice is to keep sending the full history, thinking included, and let the API drop what the current model can't use. The reasoning is gone for good only if your own client strips it.
What does account binding add?
A second wall, one that has nothing to do with the model. Blocks Sonnet 5.5 produces "work only in the account that produced them, or in an account linked to it." Send one from another account and the API drops it, again with a 200. Blocks from earlier models aren't affected.
So a conversation spread across accounts that aren't linked loses reasoning on every cross-account turn: a gateway that load-balances across several organizations' keys, rotation to a second account when the first hits a rate limit, a conversation replayed from a dev account into prod.
What do the docs leave open?
More than I'd like for a behavior that never raises an error. "Linked" isn't defined, so whether binding is per organization, workspace or key, and whether a Bedrock account counts as "another account" next to a direct API one, is not something you can read off the page. Reporting of organization_binding_mismatch is documented for the Claude API and Google Cloud only; Amazon Bedrock, Claude Platform on AWS and Microsoft Foundry aren't covered either way. The docs describe no opt-out.
Anthropic's own refusal fallback is a model switch too, and the docs say so. Server-side fallback (fallbacks: "default", beta, Claude API) retries cyber and frontier_llm refusals on Sonnet 5, which can't read Sonnet 5.5 blocks, so they're dropped. And it isn't one turn: sticky routing sends later requests in that conversation straight to the fallback model for about an hour, scoped to your organization.
How do you see the drop?
Per block, only by asking. With the thinking-binding-controls-2026-08-01 beta header, every response carries a top-level input_transformations array that lists each dropped block with its reason: model_binding_mismatch for the wrong model, organization_binding_mismatch for the wrong account. Without the header the drop is silent. A fallback does show up in the response's model field, but that tells you the model changed, not what it lost.
resp = client.beta.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
messages=history,
betas=["thinking-binding-controls-2026-08-01"],
)
for t in getattr(resp, "input_transformations", None) or []:
if t.type == "thinking_dropped":
log.warning("reasoning dropped at %s: %s", t.path, t.reason)Send your real system and tools with it, not a trimmed request. Anthropic says later checks will add new types and reasons, so never fail on one you don't recognize.
What should a router change?
The unit it routes by. My read is that the design is defensible and the default is not. A hard error would break any router that forwards thinking across an incompatible switch, and the blocks aren't destroyed, only unused for that turn. But no request fails, so ordinary error monitoring has nothing to count. The one per-block signal sits behind a beta header most integrations will never send.
A split by role, where a reviewer gets the worker's output in its own conversation, loses nothing: the reviewer reads the work product, not the worker's thinking. That's the shape of the role split I described when I put the more agentic model in the reviewer's seat. A router that picks the model turn by turn inside one conversation is now also choosing what reasoning the model gets to see. The portfolio of models I described in treating a model as a dependency that won't hold still still works, as long as each conversation stays on one member of it.
So route per conversation where you can, keep a conversation on one account, and if you do switch mid-stream, stay inside a pair that can read each other's blocks. Then turn the header on in staging, run a day of real traffic through your router, and count what it drops.