DeepSeek V4.1 Flash Live: deepseek-v4-pro Routes Away on Sep 14

Coffee Summary

  • FACT (DeepSeek, Sep 10, 2026): DeepSeek-V4.1-Flash is live on the API; set model to deepseek-flash (native multimodal).
  • FACT: deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash for compatibility; prior V4-Flash / V4-Flash-Vision-Exp are retired.
  • FACT: Starting 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches.
  • CLAIM: DeepSeek says third-party tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime — vendor-framed, not an AIImpish bench.
  • New pricing for V4.1-Flash took effect 04:00 UTC Sep 10, 2026; off-peak rates are 50% of peak (official).

What happened

DeepSeek released DeepSeek-V4.1-Flash and mirrored the notes on API docs news260910. Positioning: smallest model in a new architecture family with native visual understanding, higher throughput, and lower cache footprint versus the prior generation.

Architecture claims on the official pages: 552B-parameter MoE with a Causal Encoder–Decoder design — 8B active parameters for input, 16B for output; KV cache needing 1/4 HBM and 1/8 SSD versus previous generation (FACT: DeepSeek marketing/docs — not independently measured here).

The operationally urgent line for API customers is the Pro redirect: on Sep 14 04:00 UTC, deepseek-v4-pro traffic becomes V4.1-Flash billed at Flash rates until a future V4.1-Pro.

Why it matters

Forced model remaps break assumptions even when the new model is “better.” Latency, tool-calling quirks, multimodal behavior, and eval scores can shift under the same string you used to pin.

Your current pin Before Sep 14 04:00 UTC After cutoff (until V4.1-Pro)
`deepseek-flash` V4.1-Flash V4.1-Flash
`deepseek-v4-flash` / `…-vision-exp` Temporary → V4.1-Flash Still compatibility routes (confirm docs)
`deepseek-v4-pro` V4-Pro behavior/pricing Remapped to V4.1-Flash @ Flash rates

If you budgeted Pro unit economics or built eval gates on Pro outputs, Sep 14 is a cutover — not a soft suggestion.

What changed

1. New default Flash path — prefer deepseek-flash going forward.

2. Retirements — V4-Flash and V4-Flash-Vision-Exp retired; old IDs temporarily aliased.

3. Pro phase-out — DeepSeek is phasing out V4-Pro in favor of the Flash line until V4.1-Pro.

4. Pricing — new Flash pricing live since Sep 10 04:00 UTC; peak/off-peak continues; off-peak = 50% of peak.

5. Partners — WorkBuddy (incl. CodeBuddy) and OpenCode called out as supporting V4.1-Flash.

6. Open weights path — Hugging Face model + tech report linked from docs for self-host exploration.

Exact dollar rates are on DeepSeek’s pricing page (not invented here). Check the live pricing table before updating forecasts.

Who should care

  • Backend teams with hardcoded deepseek-v4-pro in production.
  • Agent builders sensitive to KV-cache / cache-hit billing.
  • Multimodal apps previously on Flash-Vision-Exp aliases.
  • Finance owners of DeepSeek spend comparing Pro vs Flash rate cards.
  • Teams considering self-host (HF weights) vs API.

Limitations

  • “Ahead of V4-Pro” is DeepSeek’s summary of external tests — run your own eval suite.
  • Active-parameter and KV-cache ratios are vendor architecture claims.
  • Compatibility routing duration for old Flash IDs is “temporary” — pin deepseek-flash to reduce surprise.
  • V4.1-Pro launch date is not given in the materials we used.
  • WebFetch to api-docs occasionally fails (409); content verified via curl + deepseek.com news mirror.

What to do next

Migration steps (before Sep 14 04:00 UTC)

1. Search repos/configs for deepseek-v4-pro, deepseek-v4-flash, and vision-exp aliases.

2. Stage a canary on deepseek-flash with your top 20 prompts (tools, JSON mode, vision).

3. Diff quality, latency p95, and cost vs current Pro/Flash baselines — keep a spreadsheet, not vibes.

4. Update billing alerts for Flash peak/off-peak; move batch jobs off-peak where possible.

5. If you need Pro-class behavior, plan a holdout: snapshot evals now; watch for V4.1-Pro; do not assume silent remaps preserve behavior.

6. Document the cutoff in your status page / internal changelog so support is not surprised.

7. Optional self-host: review HF DeepSeek-V4.1-Flash + tech report; large GPU+storage deployments are invited to contact DeepSeek per their post.

AIImpish Take

V4.1-Flash is both a model launch and a routing event. Treat Sep 14 like a vendor-forced migration: pin the new ID, re-run evals, and assume Pro-named traffic will behave and bill like Flash until V4.1-Pro actually exists.