Claude Opus 4.8 Fast Mode: 2.5× Speed at One-Third Prior Fast Cost

Coffee Summary

  • FACT: Anthropic launched Claude Opus 4.8 on 28 May 2026 at the same regular price as Opus 4.7: $5 / MTok input and $25 / MTok output.
  • FACT: Fast mode for Opus 4.8 runs at about 2.5× speed; Anthropic says it is three times cheaper than fast mode on prior Opus models, priced at $10 / MTok input and $50 / MTok output.
  • FACT: Effort controls landed on claude.ai (and Cowork) for all plans; Opus 4.8 defaults to high effort, with extra (xhigh in Claude Code) and max for harder work.
  • CLAIM (Anthropic evals / testers): Opus 4.8 is roughly 4× less likely than its predecessor to let code flaw

    The release pairs the model with controls that change day-to-day spend and latency — especially fast mode and effort.

    Alongside the model, Anthropic shipped effort control on claude.ai / Cowork; Dynamic Workflows in Claude Code (research preview); and Messages API mid-conversation system entries that preserve the prompt cache. This pack keeps Fast Mode, pricing, and effort in the lead; Dynamic Workflows get only a brief pointer.

    Why it matters

    For teams already on Opus, keeping the $5 / $25 regular rate lowers switching friction. The interesting economics sit in fast mode. Anthropic states Opus 4.8 fast mode works at roughly 2.5× standard generation speed and is 3× cheaper than fast mode was for previous models. Published fast-mode list prices: $10 / MTok input and $50 / MTok output — 2× regular Opus rates, but framed as about one-third prior fast cost.

    Effort control on claude.ai lets users trade depth for rate-limit headroom. Higher effort thinks more and burns limits faster; lower effort answers quicker. Opus 4.8 defaults to high, which Anthropic calls the best quality/UX balance; on coding tasks it says that level spends a similar token count to Opus 4.7’s default with better performance.

    What changed

    Use fast mode when wall-clock latency dominates. Stay on standard when batch quality-per-dollar matters more.

    Effort controls on claude.ai

    A control next to the model selector sets how hard Claude works. Anthropic recommends extra (xhigh in Claude Code) for difficult tasks and long-running async workflows, and notes higher Claude Code rate limits to absorb higher-effort token use.

    Capability / honesty signals (CLAIM)

    Anthropic and early testers report better agentic judgment and honesty — including evals that Opus 4.8 is about four times less likely than the predecessor to allow flaws in code it wrote to pass without comment. Treat partner quotes as CLAIM until you run your own suite.

    Dynamic Workflows (brief)

    Research-preview Dynamic Workflows let Claude Code plan and run large parallel subagent fleets, then verify before reporting — aimed at codebase-scale migrations (Enterprise, Team, Max). Deep coverage belongs elsewhere.

    Who should care

    API buyers comparing Opus 4.7 → 4.8 TCO; Claude Code / Max users deciding when to pay for fast mode; product managers exposing effort sliders; finance owners modeling interactive vs. batch spend.

    Limitations

    What to do next

    1. Confirm Opus 4.8 enablement and whether fast mode is unlocked on API vs. Claude Code for your tier.
    2. Price a realistic mix: % of calls on standard $5/$25 vs. fast $10/$50 at observed latency gains.
    3. Set effort defaults deliberately (high vs. extra/max) per workflow; avoid max on chatty UI paths.
    4. Re-run coding/agent evals on “did it flag its own mistakes?” — Anthropic’s honesty CLAIM is the upgrade thesis.
    5. Skim Dynamic Workflows only if you have Enterprise/Team/Max and a migration-sized job.

    AIImpish Take

    Opus 4.8 is a buyer-checklist release: same regular sticker as 4.7, a fast mode priced more like a product than a luxury tax, and effort controls on claude.ai so humans can trade depth for rate-limit headroom. Believe the 2.5× / one-third-cost story only after your own latency and invoice samples. Keep Dynamic Workflows on a separate evaluation track.

  • “2.5× speed” and “3× cheaper than prior fast” are Anthropic claims — measure tokens/sec and invoices on your traffic.
  • Partner benchmark quotes are not reproduced here.
  • Dynamic Workflows remain research preview with plan gating.
  • Mid-conversation system messages need harness discipline so instructions do not thrash caches.

Same regular price, new fast-mode math

  • Regular: $5 input / $25 output per million tokens (unchanged from 4.7).
  • Fast mode: $10 input / $50 output; ~2.5× speed; Anthropic claims ~⅓ prior fast-mode cost.
  • API model id: claude-opus-4-8.

s pass unremarked; early partners report stronger agentic judgment.

  • OPINION (AIImpish): Buy on price + fast-mode math + effort knobs first; leave Dynamic Workflows deep-dive to the dedicated pack.
  • What happened

    Anthropic upgraded the Opus line to Claude Opus 4.8, available the same day across Anthropic surfaces. The buyer headline is continuity: regular usage pricing is unchanged versus Opus 4.7.