No items found.
Start Building
Technical

DeepSeek V4 Flash API Pricing 2026: What It Actually Costs Now

DeepSeek raised API prices on August 16, 2026. Before you budget your next project, here's exactly what V4 Flash costs now — and where solo devs are finding better rates.

Author photo
packet.ai Team
August 19, 2026

DeepSeek V4 Flash costs $0.22 per million input tokens off-peak as of August 16, 2026, up from the previous $0.14 flat rate. For comparable open-weight quality at $0.59/M for Llama 3.3 70B, packet.ai Token Factory undercuts Together AI's $1.04/M rate by 43%.

Key takeaways

  • DeepSeek V4 Flash moved from a flat $0.14/M to peak/off-peak billing on August 16, 2026. Off-peak is $0.22 input / $0.66 output per million tokens.
  • Together AI still serves DeepSeek V4 Flash 0731 at $0.14 input / $0.28 output per million tokens, the old flat rate, from a third-party host.
  • packet.ai Token Factory offers Llama 3.3 70B at $0.59/M, 43% cheaper than Together AI's $1.04/M for the same model.
  • packet.ai's Llama 3.1 8B at $0.06/M is the cheapest production-grade LLM available for high-volume, cost-sensitive solo dev workloads.
  • packet.ai B200 dynamic GPU at $3.75/hr is 54% cheaper than Together AI's B200 cluster at $8.19/hr for teams self-hosting open weights.
  • Context cache hits on DeepSeek direct drop input cost to $0.007/M off-peak, a 97% discount for workloads with stable system prompts.

DeepSeek built its reputation on one thing: frontier-class model quality at prices that made Western labs uncomfortable. For most of 2026 that meant $0.14 per million input tokens, flat, simple, predictable. That changed three days ago.

On August 16, 2026, DeepSeek activated a peak and off-peak billing structure across all V4 models. For solo builders, indie hackers, and cost-sensitive teams who chose DeepSeek specifically because of the price, this matters. This guide breaks down exactly what changed, what you actually pay now across every major access route, and where packet.ai Token Factory fits into the picture.

For a broader look at LLM inference costs across all major providers, see our LLM inference cost breakdown.

What Changed on August 16, 2026: DeepSeek's New Billing Structure

Until August 16, 2026 at 16:00 UTC, DeepSeek V4 Flash billed at a flat $0.14 per million input tokens and $0.28 per million output tokens. Effective that date, DeepSeek switched to a two-tier system based on UTC time of day.

Tier Hours (UTC) Input / 1M tokens Output / 1M tokens Cache hit / 1M
V4 Flash, Off-peak All other hours $0.22 $0.66 $0.007
V4 Flash, Peak 01:00-04:00 & 06:00-10:00 $0.44 $1.32 $0.014
V4 Pro, Off-peak All other hours $0.66 $1.98 $0.022
V4 Pro, Peak 01:00-04:00 & 06:00-10:00 $1.32 $3.96 $0.044

Source: DeepSeek API documentation, accessed August 19, 2026. Always verify at api-docs.deepseek.com before committing budget.

For US-based developers, there is good news buried in the timing. Peak hours run 01:00-04:00 UTC and 06:00-10:00 UTC. That translates to roughly 8pm-1am and 1am-5am US Eastern, outside standard working hours for most solo builders running workloads during the day. If your pipeline runs on a normal US schedule, you are almost always in the off-peak window.

DeepSeek V4 Flash Cost by Provider: Direct API vs Third-Party Hosts

DeepSeek's own API is not the only way to run V4 Flash. Third-party inference providers host the same open weights, often at different prices. Here is every verified route as of August 19, 2026.

Provider Input / 1M Output / 1M Cache hit Context
OpenRouter $0.068 $0.168 $0.017 1M
Together AI $0.14 $0.28 $0.03 1M
DeepSeek direct (off-peak) $0.22 $0.66 $0.007 1M
DeepSeek direct (peak) $0.44 $1.32 $0.014 1M

Sources: together.ai/pricing (fetched August 19, 2026), openrouter.ai/deepseek/deepseek-v4-flash, api-docs.deepseek.com. Verify before use. Third-party providers may or may not mirror DeepSeek's peak/off-peak schedule.

Together AI at $0.14 input / $0.28 output is currently the cheapest route to DeepSeek V4 Flash 0731, matching the pre-hike DeepSeek direct rate. OpenRouter routes across 18 providers and shows lower blended rates depending on which provider handles your request.

packet.ai Token Factory vs Together AI: Model-for-Model Price Breakdown

For most solo dev workloads, chatbots, classification, summarisation, RAG retrieval, you do not need DeepSeek V4 specifically. You need a capable open-weight model at the lowest possible cost. packet.ai Token Factory serves several models that undercut Together AI on comparable quality tiers.

Model packet.ai Token Factory Together AI Saving Context
Llama 3.1 8B $0.06/M $0.14/M (Llama 3 8B Lite) 57% cheaper 128K
Mistral Small 3 $0.18/M Not listed n/a 32K
Llama 3.3 70B $0.59/M $1.04/M 43% cheaper 128K
Qwen 2.5 72B $0.62/M Not listed n/a 128K
DeepSeek V3 $0.85/M Not listed n/a 64K

Sources: packet.ai pricing data (August 2026), together.ai/pricing (fetched August 19, 2026). Prices per 1M tokens, blended input/output.

packet.ai Token Factory's Llama 3.3 70B at $0.59 per million tokens costs 43% less than Together AI's $1.04 rate for the same model, verified from both providers' official pricing pages on August 19, 2026.

Ready to run inference at these rates? Explore packet.ai Token Factory and start calling models via a simple OpenAI-compatible API.

Real Monthly Costs for 5 Solo Dev Workloads

Per-million-token rates are hard to apply directly. Here is what five common workloads actually cost per month across the cheapest available routes.

Workload Monthly tokens packet.ai TF Together AI Model used
Support chatbot (1K convos/day) ~20M $1.20 $2.80 Llama 3.1 8B
Document summarisation ~10M $1.80 $2.80 (V4 Flash) Mistral Small 3
RAG pipeline (50K queries/day) ~150M $9.00 $21.00 Llama 3.1 8B
Code review agent ~50M $29.50 $52.00 Llama 3.3 70B
Batch classification ~500M $30.00 $70.00 Llama 3.1 8B

Estimates assume a 60/40 input/output split. packet.ai Token Factory pricing from official spreadsheet (August 2026). Together AI pricing from together.ai/pricing (August 19, 2026).

GPU Compute for Self-Hosting: packet.ai B200 vs Together AI

If your volume crosses roughly 5 million output tokens per day, self-hosting DeepSeek V4 open weights on your own GPU starts to beat the API. DeepSeek released V4 Flash under an MIT licence, so the weights are yours to run.

GPU packet.ai / GPU / hr Together AI / GPU / hr packet.ai saving
B200 (dynamic / on-demand) $3.75 $8.19 54% cheaper
B200 (dedicated) $5.90 $8.99 (dedicated) 34% cheaper
H100 (on-demand) Launching Soon $3.99 n/a
H200 (on-demand) Launching Soon $5.99 n/a

Sources: packet.ai pricing (August 2026), together.ai/pricing (fetched August 19, 2026). All prices per GPU per hour.

packet.ai B200 dynamic at $3.75 per GPU per hour is 54% cheaper than Together AI's on-demand B200 cluster at $8.19 per hour, making it the most cost-effective path for teams running DeepSeek V4 open weights at scale.

To deploy a B200 cluster for self-hosted inference, see packet.ai B200 GPU pricing and availability. For a broader set of GPU options, browse available GPU clusters.

How to Call DeepSeek V4 Flash API: 3-Line Setup

DeepSeek V4 Flash uses an OpenAI-compatible endpoint. If you have an existing OpenAI SDK integration, migrating is a base URL change and a model name update.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DEEPSEEK_API_KEY",
    base_url="https://api.deepseek.com"
)

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)

Migration required

The legacy deepseek-chat and deepseek-reasoner model aliases were retired on July 24, 2026. Any code still calling these names will error. Update to deepseek-v4-flash (non-thinking and thinking modes) or deepseek-v4-pro.

Context Caching: The 97% Discount Most Developers Miss

DeepSeek applies disk-based prefix caching automatically. No SDK changes or special headers needed. When your prompt shares the same prefix across requests (a system prompt, a retrieved document, tool definitions), the cached portion bills at the cache-hit rate instead of the standard input rate.

At off-peak rates, a cache hit costs $0.007 per million tokens versus $0.22 for a cache miss, a 97% reduction on that portion of your input. For agent loops with a stable instruction set or RAG pipelines with a consistent document prefix, this can cut your effective input cost to near zero.

The practical rule: put stable content (instructions, schemas, documents) at the start of your prompt and keep it byte-identical across calls. Variable content (the user's question) goes at the end.

Which Model and Provider Should You Use? A Solo Dev Decision Guide

The right choice depends on your daily volume and what quality floor you need. Here is a direct decision matrix based on verified pricing.

Daily output volume Best route Approx. monthly cost
Under 1M tokens packet.ai Token Factory, Llama 3.1 8B ($0.06/M) Under $2/mo
1M-10M tokens packet.ai Token Factory, Mistral Small 3 ($0.18/M) or Llama 3.3 70B ($0.59/M) $5-$180/mo
10M-100M tokens Together AI, DeepSeek V4 Flash ($0.14/$0.28) or packet.ai Token Factory $42-$420/mo
100M+ tokens Self-host on packet.ai B200 ($3.75/hr dynamic) ~$2,700/mo fixed

For a deeper look at how self-hosting costs compare to managed inference APIs across different GPU types, see our AI token cost guide.

Frequently Asked Questions

As of August 16, 2026, DeepSeek V4 Flash bills at $0.22 input / $0.66 output per million tokens off-peak, and $0.44 input / $1.32 output at peak (01:00-04:00 and 06:00-10:00 UTC). Cache hits cost $0.007 per million tokens off-peak. Third-party providers like Together AI still offer DeepSeek V4 Flash 0731 at the old $0.14/$0.28 flat rate.
Yes for several key models. packet.ai Token Factory serves Llama 3.3 70B at $0.59/M, 43% cheaper than Together AI's $1.04/M for the same model. Llama 3.1 8B costs $0.06/M on packet.ai, versus $0.14/M on Together AI for a comparable 8B model. Pricing verified from both providers' official sources on August 19, 2026.
DeepSeek peak hours run 01:00-04:00 UTC and 06:00-10:00 UTC. In US Eastern time, that is 8pm-11pm and 1am-5am, outside normal working hours. Developers running pipelines during US business hours (9am-6pm ET) are almost always in the off-peak window at $0.22/$0.66 per million tokens, not the peak rate of $0.44/$1.32.
Self-hosting becomes cost-effective at roughly 5 million output tokens per day. Below that threshold, packet.ai Token Factory or Together AI serverless is cheaper. Above it, a single B200 GPU on packet.ai at $3.75/hr, 54% cheaper than Together AI's $8.19/hr, starts paying for itself versus per-token API billing. DeepSeek V4 Flash weights are MIT-licensed and publicly available.
Yes. DeepSeek uses an OpenAI-compatible endpoint at api.deepseek.com. Set your base_url to that address, swap in your DeepSeek API key, and update the model parameter to deepseek-v4-flash. No other code changes are needed. Note that the legacy deepseek-chat and deepseek-reasoner aliases were retired on July 24, 2026. Update any code still using those names or requests will error.

Last reviewed: August 19, 2026. For the most affordable hosted LLM inference, explore packet.ai Token Factory, or browse GPU clusters if you are ready to self-host open weights.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog