Claude Sonnet 5 Migration: Adaptive Thinking Defaults and Sampling 400s

Coffee Summary

  • FACT: Claude Sonnet 5 (claude-sonnet-5) is a drop-in upgrade from Sonnet 4.6 at $2 / $10 per million input/output tokens (vs 4.6’s $3 / $15), with 1M context and 128k max output.
  • FACT (breaking): Adaptive thinking is on by default; thinking: {type: "enabled", budget_tokens: N} returns 400. Use adaptive + What happened

    Anthropic’s Sonnet 5 docs position the model as the speed/intelligence sweet spot — overview release date 30 Jun 2026 — and a drop-in upgrade from Claude Sonnet 4.6 with a short list of hard API breaks. If you only change the model ID and keep old thinking/sampling knobs, production traffic can start failing with 400s.

    Official path: update claude-sonnet-4-6 → claude-sonnet-5, remove manual extended thinking and non-default sampling parameters, fix parsers that assume content[0] i

    What changed

    Adaptive thinking default (FACT — breaking)

    On Sonnet 4.6, omitting thinking meant no thinking. On Sonnet 5, those requests run adaptive thinking. Manual extended thinking (type: "enabled" + budget_tokens) returns 400.

    Use adaptive thinking and steer with output_config.effort, or opt out with thinking: {type: "disabled"}. Default thinking.display is omitted (often summarized on 4.6); set display: "summarized" if UIs show thinking text. Responses may start with thinking blocks before text — select by type, and pass thinking blocks back unmodified in tool loops.

    Sampling parameters → 400 (FACT — breaking)

    Non-default temperature, top_p, or top_k return 400. Omit them. Guide behavior with prompts or structured outputs.

    s text, and re-run token counting. Claude Code can automate much of this with /claude-api migrate.

    Why it matters

    Sonnet-class traffic is usually the workhorse lane. Silent behavior change (thinking now on) plus hard rejects on legacy params means latency, bill shape, and error rates move together. Teams that used temperature as a creativity dial or budget_tokens as a cost governor need a new control: effort (default high).

    The tokenizer shift is the quiet budget issue: fewer dollars per token, more tokens per string — so a lower $/MTok is not a 1:1 cheaper endpoint.

    e>effort, or thinking: {type: "disabled"}.

  • FACT (breaking): Non-default temperature, top_p, or top_k return 400.
  • FACT: New tokenizer ≈ ~30% more tokens for the same text — recount usage; equivalent-request cost does not fall 1:1 with lower per-token pricing.
  • FACT: On Claude API and Google Cloud, Sonnet 5 adds stable computer_toolset_20260801 and browser use; Priority Tier is not available on Sonnet 5.