DeepSeek V4.1 Flash Live: deepseek-v4-pro Routes Away on Sep 14
Coffee Summary
- FACT (DeepSeek, Sep 10, 2026): DeepSeek-V4.1-Flash is live on the API; set model to
deepseek-flash(native multimodal). - FACT:
deepseek-v4-flashanddeepseek-v4-flash-vision-exptemporarily route to V4.1-Flash for compatibility; prior V4-Flash / V4-Flash-Vision-Exp are retired. - FACT: Starting 04:00 UTC on September 14, 2026, all
deepseek-v4-prorequests route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches. - CLAIM: DeepSeek says third-party tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime — vendor-framed, not an AIImpish bench.
- New pricing for V4.1-Flash took effect 04:00 UTC Sep 10, 2026; off-peak rates are 50% of peak (official).
What happened
DeepSeek released DeepSeek-V4.1-Flash and mirrored the notes on API docs news260910. Positioning: smallest model in a new architecture family with native visual understanding, higher throughput, and lower cache footprint versus the prior generation.
Architecture claims on the official pages: 552B-parameter MoE with a Causal Encoder–Decoder design — 8B active parameters for input, 16B for output; KV cache needing 1/4 HBM and 1/8 SSD versus previous generation (FACT: DeepSeek marketing/docs — not independently measured here).
The operationally urgent line for API customers is the Pro redirect: on Sep 14 04:00 UTC, deepseek-v4-pro traffic becomes V4.1-Flash billed at Flash rates until a future V4.1-Pro.
Why it matters
Forced model remaps break assumptions even when the new model is “better.” Latency, tool-calling quirks, multimodal behavior, and eval scores can shift under the same string you used to pin.
| Your current pin | Before Sep 14 04:00 UTC | After cutoff (until V4.1-Pro) |
|---|---|---|
| `deepseek-flash` | V4.1-Flash | V4.1-Flash |
| `deepseek-v4-flash` / `…-vision-exp` | Temporary → V4.1-Flash | Still compatibility routes (confirm docs) |
| `deepseek-v4-pro` | V4-Pro behavior/pricing | Remapped to V4.1-Flash @ Flash rates |
If you budgeted Pro unit economics or built eval gates on Pro outputs, Sep 14 is a cutover — not a soft suggestion.
What changed
1. New default Flash path — prefer deepseek-flash going forward.
2. Retirements — V4-Flash and V4-Flash-Vision-Exp retired; old IDs temporarily aliased.
3. Pro phase-out — DeepSeek is phasing out V4-Pro in favor of the Flash line until V4.1-Pro.
4. Pricing — new Flash pricing live since Sep 10 04:00 UTC; peak/off-peak continues; off-peak = 50% of peak.
5. Partners — WorkBuddy (incl. CodeBuddy) and OpenCode called out as supporting V4.1-Flash.
6. Open weights path — Hugging Face model + tech report linked from docs for self-host exploration.
Exact dollar rates are on DeepSeek’s pricing page (not invented here). Check the live pricing table before updating forecasts.
Who should care
- Backend teams with hardcoded
deepseek-v4-proin production. - Agent builders sensitive to KV-cache / cache-hit billing.
- Multimodal apps previously on Flash-Vision-Exp aliases.
- Finance owners of DeepSeek spend comparing Pro vs Flash rate cards.
- Teams considering self-host (HF weights) vs API.
Limitations
- “Ahead of V4-Pro” is DeepSeek’s summary of external tests — run your own eval suite.
- Active-parameter and KV-cache ratios are vendor architecture claims.
- Compatibility routing duration for old Flash IDs is “temporary” — pin
deepseek-flashto reduce surprise. - V4.1-Pro launch date is not given in the materials we used.
- WebFetch to api-docs occasionally fails (409); content verified via curl + deepseek.com news mirror.
What to do next
Migration steps (before Sep 14 04:00 UTC)
1. Search repos/configs for deepseek-v4-pro, deepseek-v4-flash, and vision-exp aliases.
2. Stage a canary on deepseek-flash with your top 20 prompts (tools, JSON mode, vision).
3. Diff quality, latency p95, and cost vs current Pro/Flash baselines — keep a spreadsheet, not vibes.
4. Update billing alerts for Flash peak/off-peak; move batch jobs off-peak where possible.
5. If you need Pro-class behavior, plan a holdout: snapshot evals now; watch for V4.1-Pro; do not assume silent remaps preserve behavior.
6. Document the cutoff in your status page / internal changelog so support is not surprised.
7. Optional self-host: review HF DeepSeek-V4.1-Flash + tech report; large GPU+storage deployments are invited to contact DeepSeek per their post.
AIImpish Take
V4.1-Flash is both a model launch and a routing event. Treat Sep 14 like a vendor-forced migration: pin the new ID, re-run evals, and assume Pro-named traffic will behave and bill like Flash until V4.1-Pro actually exists.
AIImpish