Start Building →
Guide

DeepSeek V4.1 Flash: The New Model, GPU Sizing, and API Guide

DeepSeek retired its old Flash model and reversed a plan to retire Pro too. Here is current pricing, plus which GPUs can actually run it.

Author photo
packet.ai Team
September 28, 2026

DeepSeek new model releases have a habit of reshuffling the pricing conversation, and the latest one is no exception. DeepSeek's newest and latest model, V4.1 Flash, shipped September 10, 2026, replacing the previous Flash tier entirely, adding native image understanding, and beating DeepSeek's own Pro tier on several benchmarks at a fraction of the cost, though the comparison is genuinely mixed across benchmark families. This guide covers what actually changed with deepseek v4 1, current pricing and specs, and what GPUs can actually run it given the model's genuinely large deployable footprint.

Key takeaways

  • DeepSeek V4.1 Flash is a genuinely new architecture, not a tuned version of the prior Flash model: a 552B-backbone-parameter mixture-of-experts design activating roughly 8B parameters on input and 16B on output, with native image understanding built in rather than served as a separate experimental model
  • DeepSeek retired the old Flash model outright. The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp model IDs still work, but every request through them is now served by V4.1 Flash and billed at Flash pricing; the current model ID going forward is deepseek-flash
  • Official DeepSeek pricing is $0.15 input and $0.60 output per million tokens off-peak, doubling to $0.30 and $1.20 during peak hours (01:00-04:00 and 06:00-10:00 UTC on weekdays), with cache-hit input priced separately at $0.003 off-peak and $0.006 peak
  • DeepSeek initially planned to route its Pro tier to V4.1 Flash entirely starting September 14, 2026, then reversed that decision days later; V4 Pro remains available as a separate model rather than being fully retired
  • The 552B backbone figure understates the real deployable footprint: DeepSeek's own model card and independent vLLM serving recipes put the actual checkpoint at roughly 511 GB on disk, with a 614 GB minimum VRAM budget once serving overhead is included, and validated multi-GPU configurations on H200, GB200, GB300, and MI350X hardware

What Actually Changed With V4.1 Flash

DeepSeek V4.1 Flash is built on a new Causal Encoder-Decoder architecture, a genuine departure from the prior Flash model's design rather than an incremental deepseek update. The model has 552B backbone parameters as a mixture-of-experts design, activating roughly 8B parameters during input processing and 16B during output generation, an asymmetric activation pattern that's itself a departure from the previous Flash tier's more conventional decoder-only setup. The architecture separates input processing from generation: the decoder's KV cache is projected from the encoder's final hidden states rather than computed fresh at every decoder layer, which is specifically what allows the lower per-token activation during prefill.

The most immediately visible change is native image understanding, making V4.1 Flash a genuine deepseek vision model and deepseek multimodal release rather than a text-only update. Previously, DeepSeek offered vision capability through a separate experimental model, deepseek-v4-flash-vision-exp, alongside the text-only mainline Flash model. V4.1 Flash folds that capability directly into the primary model rather than keeping it as a separate experimental branch, and DeepSeek's own official pricing page confirms vision support as a standard feature of the current deepseek-flash model, not an add-on.

DeepSeek's own published evaluations show a genuinely mixed benchmark picture against the Pro tier it's positioned beneath, not a simple win across the board. V4.1 Flash leads on several coding and reasoning benchmarks, including BigCodeBench, HumanEval, and GSM8K, while V4 Pro remains ahead on others, including AGIEval, SimpleQA-Verified, SuperGPQA, and MATH. The honest summary is that V4.1 Flash closes much of the gap to Pro at a fraction of the cost, rather than uniformly beating it.

Model IDs, Retirement, and the V4 Pro Reversal

DeepSeek retired the previous Flash model outright rather than running it alongside the new one. The legacy model IDs, deepseek-v4-flash and deepseek-v4-flash-vision-exp, are still accepted by the API for backward compatibility, but every request sent to either one is now actually served by DeepSeek-V4.1-Flash and billed at V4.1 Flash's pricing. The current, forward-looking model ID is deepseek-flash. Existing integrations using the old ID strings can continue working, but those IDs now resolve to V4.1 Flash rather than the retired V4 Flash model.

A more unusual development involves DeepSeek's Pro tier. DeepSeek initially announced that starting September 14, 2026, every request to deepseek-v4-pro would be routed to V4.1 Flash and billed at Flash pricing, effectively retiring Pro as a distinct model. That plan was reversed before it took effect: DeepSeek-V4-Pro-0813 remains available as its own model on the current pricing page, with its own separate pricing tier, rather than being folded into Flash. Anyone who built integrations around the originally announced routing change should verify current behavior directly, since the plan that was publicly announced is not the plan that actually shipped.

⚡ Current official deepseek flash pricing, direct from DeepSeek (last checked September 29, 2026)

DeepSeek's own pricing page lists deepseek-flash at $0.15 per million input tokens and $0.60 per million output tokens off-peak, doubling to $0.30 and $1.20 during peak hours. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; all other hours, including weekends, are billed at the off-peak rate. Cache-hit input tokens are priced separately and far lower, at $0.003 off-peak and $0.006 peak. Third-party resellers and API proxies often list different, higher figures that include their own markup on top of these rates, so verify against DeepSeek's own pricing page directly rather than a reseller's listing if exact cost matters.

Deepseek flash api: Specs, Context Window, and Features

DeepSeek V4.1 Flash supports a deepseek flash context window of 1M tokens with a maximum output of 384K tokens, matching the context capacity of the Pro tier despite the significant cost difference between the two. The deepseek flash api supports both thinking and non-thinking modes, switchable per request, JSON output mode, tool calling, and both OpenAI-format and Anthropic-format requests, using protocol-specific base URLs: https://api.deepseek.com for OpenAI-format requests, and https://api.deepseek.com/anthropic for Anthropic-format requests.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DEEPSEEK_API_KEY",
    base_url="https://api.deepseek.com"
)

response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{"role": "user", "content": "Hello"}]
)

The documented concurrency limit is 2,500 requests for deepseek-flash, compared with 500 for deepseek-v4-pro. That matters when estimating how much parallel API traffic a production integration can send through DeepSeek's own infrastructure directly, separate from any self-hosted deployment's own capacity.

What GPUs Can Actually Run DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is not a single-GPU model despite its relatively low active-parameter count. The 552B backbone figure also understates the real deployable footprint: DeepSeek's model card separately describes additional components, including roughly 196B parameters of Engram conditional memory, on top of the routed-expert backbone. Independent vLLM serving recipes put the full checkpoint at approximately 511 GB on disk (476 GiB), broken down roughly as 259.5 GiB for routed and speculative-decoding experts, 188.8 GiB for the Engram tables, and the remainder for attention, embedding, and quantization scale components.

HardwareVerified configurationNotes
H200Single 8-GPU nodeVerified in the official vLLM recipe; room for KV cache, though very long contexts may need capping
GB200 NVL4TP4, one trayDefault hardware target in the vLLM recipe; also verified in a disaggregated prefill/decode setup
GB300VerifiedListed as verified hardware in the official recipe
MI350X (AMD)4 of 8 GPUs on a nodeVerified with AMD-specific serving backend flags

These are current, independently documented serving configurations, not universal minimums for every workload; the exact requirement still depends on weight format, parallelism strategy, KV cache settings, and context length actually used in production.

What Self-Hosting V4.1 Flash Actually Requires

The mixture-of-experts architecture's active-parameter count, roughly 8B on input and 16B on output, describes compute per token, not deepseek flash vram or memory footprint. Self-hosting requires enough aggregate memory to load the model's full deployable checkpoint, not just the parameters activated for each token; MoE architectures route each token through a subset of experts dynamically, but the full set of experts still has to be resident and addressable. This is the same distinction that applies to every large MoE model: modest active compute doesn't translate to modest storage or memory requirements. On any deepseek flash benchmark measuring cost per intelligence, V4.1 Flash's positioning as deepseek cheapest model in the current lineup only holds on the managed API; self-hosting removes that specific pricing advantage and replaces it with infrastructure cost instead.

The open weights are published under the MIT license, so self-hosting DeepSeek's latest model, deepseek latest model releases having followed the same MIT pattern for several generations now, is genuinely available as an option, not restricted the way some open-weight releases are. Whether self-hosting or using the managed API makes more sense for a given workload comes down to the same utilization math that applies to any large open-weight model: infrastructure and engineering overhead against API pricing, scaled against actual request volume. For a closer look at how active parameters relate to actual VRAM needs specifically for DeepSeek's model family, the packet.ai DeepSeek GPU requirements guide covers the V4 generation in more depth.

Running DeepSeek V4.1 Flash Through a Managed Endpoint

Given the model's genuinely large deployable footprint despite its modest active-parameter count, a managed endpoint removes the specific infrastructure burden of provisioning and operating that GPU cluster while still providing access to the model itself. This is a different option than DeepSeek's own direct API covered above: it's a managed endpoint for the underlying open-weight model, not DeepSeek's own hosted service. packet.ai's Token Factory is being built as an OpenAI-compatible endpoint for open-weight models, which could put a model like V4.1 Flash behind the same interface already covered in the packet.ai OpenAI-compatible migration guide, and more broadly in the packet.ai What Is an LLM API guide. It's currently in private preview, with the specific model catalog and pricing still being finalized; V4.1 Flash's inclusion is not confirmed until that catalog is published.

Join the waitlist for early access once it opens.

Sources and Further Reading

Frequently asked questions

If you're asking what is deepseek flash: DeepSeek V4.1 Flash, released September 10, 2026, is DeepSeek's newest model, replacing the previous Flash tier entirely. It's a 552B-backbone-parameter mixture-of-experts model with native image understanding, a 1M-token context window, 8B parameters activated during prefill and 16B during decode, and pricing significantly below DeepSeek's Pro tier, callable via the deepseek-flash model ID.
Use deepseek-flash for new API integrations. The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs remain accepted for backward compatibility, but both now route to and are billed as V4.1 Flash.
Per DeepSeek's official pricing page: $0.15 input and $0.60 output per million tokens off-peak, doubling to $0.30 and $1.20 during peak hours (01:00-04:00 and 06:00-10:00 UTC, weekdays). Cache-hit input is priced separately at $0.003 off-peak and $0.006 peak. Third-party resellers often list higher, marked-up figures, so verify against DeepSeek's own pricing page for exact cost.
There's no single universal number, but this is a multi-GPU model. Independent serving recipes put the checkpoint at roughly 476 GiB on disk with a 614 GB minimum VRAM budget, and validate configurations including a single 8-GPU H200 node and a 4-GPU GB200 NVL4 tray. Do not size a deployment from the 8B/16B active-parameter count alone.

Last reviewed: September 29, 2026. DeepSeek's pricing, model routing, and API behavior have changed multiple times since V4.1 Flash's September 10, 2026 launch, including a reversed routing decision for V4 Pro; verify current model IDs and pricing directly against DeepSeek's official pricing page before building on this model.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog