Start Building
Alternative

Cheapest LLM API Providers in 2026: Price Per Million Tokens

DeepSeek repriced on August 16. GPT-5.6 Luna dropped 80% on July 30. The cheapest LLM API list changes faster than most teams track it. Here is where every major provider sits right now.

Author photo
packet.ai Team
August 24, 2026

The cheapest LLM API in August 2026 is Groq's Llama 3.1 8B at $0.05 per million input tokens - but the cheapest capable model at production quality is DeepSeek V4 Flash at $0.22/$0.66 per million tokens off-peak after the August 16 repricing, with packet.ai Token Factory serving Llama 3.3 70B at $0.59/M as the lowest rate for that model class.

Key takeaways

  • LLM API prices dropped approximately 80% between early 2025 and August 2026. GPT-4o input fell from $5.00 to $0.20 per million tokens across the same model family within 18 months.
  • DeepSeek V4 Flash repriced on August 16, 2026, moving from a flat $0.14/$0.28 to peak/off-peak billing at $0.22/$0.66 off-peak and $0.44/$1.32 during peak hours (01:00-04:00 and 06:00-10:00 UTC).
  • Output tokens cost 3 to 6x more than input tokens across every major provider. For response-heavy workloads, the output multiplier dominates the bill more than the input rate.
  • packet.ai Token Factory serves Llama 3.3 70B at $0.59/M - 43% below Together AI's $1.04/M for the same model - with no cold starts and no idle billing.
  • Batch API discounts (50% off at OpenAI, Anthropic, and Google) and prompt caching (cache hits at 0.1x input price at Anthropic and DeepSeek) are the two levers that halve a real bill without changing the model.
  • Free tiers from Groq (30 RPM, no credit card), OpenRouter, and Google AI Studio cover most prototyping workloads. Paid APIs become necessary when you exceed rate limits or need guaranteed throughput.

Two things happened to LLM API pricing in the six weeks before this post. On July 30, 2026, OpenAI cut GPT-5.6 Luna by 80%, dropping it to $0.20 per million input tokens inside a current flagship family. On August 16, DeepSeek moved V4 Flash from a flat rate to peak/off-peak billing, effectively raising costs during business hours by 50 to 200% depending on the line item. Budgets written in July are wrong in August. This cheapest llm api comparison covers every major provider as of August 24, 2026 with rates verified against official pricing pages and cross-checked against independent tracking services. For the self-hosting math behind these numbers, see the LLM inference cost breakdown.

How LLM API Token Pricing Works: Input, Output, and the Multiplier That Matters

Every major LLM API bills the same way: per token, split into input and output, quoted per million tokens (MTok). A token is roughly 0.75 words in English, so 1 million tokens is approximately 750,000 words - about 1,500 pages of text. Input tokens cover everything you send: your prompt, system instructions, conversation history, and any documents in context. Output tokens cover what the model generates back.

Output tokens cost 3 to 6x more than input tokens at every major provider - and that multiplier matters more than the headline input rate for most workloads. GPT-5.6 Sol charges 6x its input rate for output. Claude Sonnet 5 charges 5x. DeepSeek V4 Flash charges 3x off-peak. A workload that sends 2,000-token prompts and receives 500-token responses has 80% of its token volume on the cheap side. A chatbot that generates 2,000-token responses to 200-token questions has the opposite problem: the output multiplier dominates the bill.

80%

price drop since early 2025

3-6x

output vs input token cost

50%

batch API discount (OpenAI, Anthropic, Google)

$0.05

cheapest production API input (Groq Llama 3.1 8B)

Three discounts change effective cost without changing the model. Batch APIs at OpenAI, Anthropic, and Google cut every rate 50% for async workloads with up to 24-hour turnaround. Prompt caching drops repeated input prefixes - system prompts, document chunks sent across many requests - to 0.1x the base input price at Anthropic and a similar rate at DeepSeek. Context caching at Google Gemini cuts 80-90% off standard rates for cached content. Stack batch and caching on Anthropic and you pay roughly 25% of on-demand for cached, async workloads.

DeepSeek API Pricing in August 2026: V4 Flash and V4 Pro After the Repricing

DeepSeek is the most-searched LLM API provider in August 2026, with search volume up 662% year-over-year on "deepseek api pricing" - and for good reason. V4 Flash was the cheapest capable LLM API on the market at $0.14/$0.28 for most of the year. That changed on August 16, 2026.

On August 16, 2026 at 16:00 UTC, DeepSeek activated peak/off-peak billing across all V4 models. The old flat rate of $0.14/$0.28 for V4 Flash is gone from DeepSeek's own API. Off-peak rates (outside 01:00-04:00 and 06:00-10:00 UTC) are $0.22 input / $0.66 output per million tokens. During peak windows, rates jump to $0.44 input / $1.32 output. For teams running workloads timed to US business hours, the effective rate is the peak rate - not the off-peak headline figure. For a detailed breakdown of what changed and how to route around peak pricing, see the DeepSeek V4 Flash pricing guide.

Model Input /M Output /M Cache hit input Context window Note
V4 Flash (off-peak) $0.22 $0.66 $0.007 1M tokens Outside 01-04 and 06-10 UTC
V4 Flash (peak) $0.44 $1.32 $0.014 1M tokens 01:00-04:00 and 06:00-10:00 UTC
V4 Pro (off-peak) $0.66 $1.98 $0.0036 1M tokens Reasoning model tier
V4 Pro (peak) $1.32 $3.96 $0.0072 1M tokens 01:00-04:00 and 06:00-10:00 UTC

Source: DeepSeek API documentation, accessed August 21, 2026 via morphllm.com. Third-party providers including Together AI still serve DeepSeek V4 Flash 0731 at the pre-August 16 flat rate of $0.14/$0.28 per million tokens. If flat-rate billing matters for your budget predictability, those third-party hosts are currently the cheapest access route for V4 Flash.

DeepSeek V4 Flash cache hits cost $0.007 per million input tokens off-peak - roughly 97% below the on-demand input rate. For RAG pipelines that send the same system prompt or document context across hundreds of requests, caching is the single largest cost lever available on the DeepSeek API.

LLM API Pricing Comparison: Every Major Provider Ranked by Cost, August 2026

All rates below are on-demand per million tokens, verified August 24, 2026. Batch API and caching discounts are noted separately - apply them to on-demand figures to get your real effective rate.

Provider / Model Input /M Output /M Batch discount Free tier Best for
packet.ai Token Factory
Llama 3.3 70B
$0.59 $0.59 No 1M tokens/day Open-model inference, no cold starts
packet.ai Token Factory
Llama 3.1 8B
$0.06 $0.06 No 1M tokens/day Budget inference, classification, tagging
Groq
Llama 3.1 8B Instant
$0.05 $0.08 50% 30 RPM Lowest absolute rate, LPU speed
Gemini 2.5 Flash-Lite
Google
$0.10 $0.40 50% Yes Google ecosystem, context caching
Llama 4 Maverick
Together AI / Fireworks
$0.15 $0.60 Varies Limited Open-weight frontier quality at commodity price
DeepSeek V4 Flash
Off-peak (DeepSeek direct)
$0.22 $0.66 No Limited High-quality coding and reasoning at low cost
DeepSeek V4 Flash
Together AI (flat rate)
$0.14 $0.28 Varies Limited Flat-rate DeepSeek V4 Flash 0731 checkpoint
GPT-5.6 Luna
OpenAI
$0.20 $1.20 50% Limited Frontier family at budget pricing (post July 30 cut)
Mistral Small
Mistral AI
$0.20 $0.60 Varies Limited EU data residency, low output cost
Claude Haiku 4.5
Anthropic
$1.00 $5.00 50% No Anthropic quality at budget Anthropic pricing
Claude Sonnet 5
Anthropic (until Aug 31)
$2.00 $10.00 50% No Reverts to $3/$15 on September 1
Gemini 3.1 Pro
Google
$2.00 $12.00 50% Yes Doubles above 200K tokens ($4/$18)
Claude Opus 5
Anthropic
$5.00 $25.00 50% No Frontier quality ceiling, coding and reasoning

All rates are on-demand per million tokens. Batch API and caching discounts apply on top. Sources: DeepSeek API docs (Aug 21, 2026), morphllm.com, IntuitionLabs, official provider pricing pages. Verify before committing.

packet.ai Token Factory Llama 3.3 70B at $0.59/M is 43% below Together AI's $1.04/M for the same model, with no cold starts, no idle billing, and an OpenAI-compatible API. For teams running GPU workloads on packet.ai already, Token Factory tokens and GPU-hour credits share the same billing account.

packet.ai Token Factory: Cheapest AI API for Open Models at Production Scale

Token Factory is packet.ai's managed inference API: OpenAI-compatible endpoint, pay-per-token, no GPU management. It runs on packet.ai's GPU infrastructure, which means the margin structure is different from third-party resellers. Llama 3.3 70B at $0.59/M input and $0.59/M output is the same rate for both directions - no output premium - because the GPU cost at packet.ai's H200 and B200 rates makes that math work.

The current model catalog includes Llama 3.1 8B at $0.06/M, Llama 3.3 70B at $0.59/M, DeepSeek models, Qwen, and Kimi K3 at $1.50/$7.50 per million tokens - 50% below Together AI's list price for Kimi K3. The free tier gives 1 million tokens per day during development with no credit card required. For teams scaling past the free tier, there is no minimum commitment and no contract.

The crossover where self-hosting on packet.ai GPUs beats the Token Factory API is roughly 12 million output tokens per month per model, as the LLM inference cost breakdown shows in detail. Below that threshold, Token Factory is cheaper than the GPU rental math.

Best for: Teams that need OpenAI-compatible inference on open models without managing vLLM, TGI, or SGLang infrastructure. Particularly cost-effective for Llama 3.3 70B and Kimi K3 versus third-party hosts. Explore Token Factory pricing and available models.

Cheapest AI API by Workload Tier: Budget, Mid-Range, and Frontier

The right cheapest LLM API depends entirely on which quality tier your workload needs. A classification pipeline running 100 million tokens per day on Llama 3.1 8B at $0.05/M pays $5/day. The same workload on Claude Opus 5 at $5/M input pays $500/day - 100x more for quality overkill. Most production systems route by task: cheap small models for classification and extraction, mid-tier for summarisation and drafting, frontier only for the subset of queries that require it.

01

Budget tier: $0.05 - $0.22 per million input tokens

Groq Llama 3.1 8B at $0.05/M, packet.ai Token Factory Llama 3.1 8B at $0.06/M, Gemini 2.5 Flash-Lite at $0.10/M, Llama 4 Maverick on Together/Fireworks at $0.15/M. Right for: classification, tagging, embedding pipelines, structured extraction, high-volume batch workloads where response quality is not the primary constraint.

02

Mid-range tier: $0.20 - $2.00 per million input tokens

GPT-5.6 Luna at $0.20/M, Mistral Small at $0.20/M, DeepSeek V4 Flash off-peak at $0.22/M, packet.ai Token Factory Llama 3.3 70B at $0.59/M, Kimi K3 on Token Factory at $1.50/M, Claude Sonnet 5 at $2.00/M (until Aug 31). Right for: chatbots, summarisation, code generation, drafting, RAG applications with moderate quality requirements.

03

Frontier tier: $5.00+ per million input tokens

Claude Opus 5 at $5/M, GPT-5.6 Sol at $5/M. The output cost is where it compounds: Opus 5 at $25/M output, Sol at $30/M. Right for: complex reasoning, agentic workflows, code review, tasks where the quality gap materially affects outcomes and volume is low enough that cost is secondary. At 1 million output tokens, the difference between Haiku 4.5 ($5) and Opus 5 ($25) is $20. At 100 million output tokens, it is $2 million.

The most cost-effective production architecture is not picking one model - it is routing. Send classification and entity extraction to Llama 3.1 8B at $0.06/M. Send summarisation and drafting to Llama 3.3 70B at $0.59/M. Reserve DeepSeek V4 Flash or GPT-5.6 Luna for complex generation. Reserve Opus 5 or Kimi K3 for the 5-10% of queries that need frontier reasoning. Morph's model router, OpenRouter, and LiteLLM all support this pattern without code changes per model.

Free LLM API Tiers in 2026: Where to Start Without Paying

For most prototyping workloads, the cheapest LLM API is free. Three providers offer genuinely usable free tiers in August 2026. Groq's free tier runs all supported models (Llama, Qwen, Kimi K2) at 30 requests per minute and 14,400 requests per day with no credit card required - fast enough for development and benchmarking on smaller workloads. Google AI Studio provides free Gemini access with generous rate limits for individuals. packet.ai Token Factory provides 1 million tokens per day free during development.

The practical constraint on free tiers is rate limits, not cost. Groq's 30 RPM ceiling means 1,800 requests per hour - adequate for prototyping a chatbot or running evaluation sweeps, insufficient for production inference above a few hundred concurrent users. When free tier rate limits become the constraint, the decision is whether to pay $0.05/M on Groq's paid tier or move to a provider with higher throughput guarantees at a similar price point.

OpenRouter is worth knowing for prototyping: it routes each request to the cheapest stable provider automatically, at provider list price with no markup, and provides a single API key that works across all connected providers. It is not a production architecture - latency variability and provider routing changes are real - but for evaluation and cost benchmarking across providers, it eliminates the multi-account setup problem.

Self-Hosted LLM vs Managed API: When the Crossover Happens

Self-hosting DeepSeek V3 or Llama 4 on packet.ai GPUs becomes cheaper than the managed API above a specific token volume threshold. The math is straightforward: an 8-GPU H200 node on packet.ai at $2.49/hr costs approximately $1,800/month running continuously. That node serves the full 671B DeepSeek V3 open-weight model at FP8 quantization. The break-even against DeepSeek V4 Flash at $0.22/M input and $0.66/M output off-peak is roughly 2-3 billion tokens per month at typical prompt/completion ratios - a threshold most teams do not hit.

For smaller open-weight models (7B to 70B), the math shifts. A single H100 PCIe on packet.ai at $2.50/hr serves Llama 3.3 70B at approximately 3,000 tokens per second with vLLM and PagedAttention enabled. At 85% utilisation, that node generates roughly 220 million tokens per day. Against packet.ai Token Factory's $0.59/M rate, the crossover is approximately 12 million output tokens per month - reachable for a moderately busy production chatbot. Below that, Token Factory is cheaper. Above it, your own GPU is cheaper.

For teams evaluating the GPU rental side of this decision, packet.ai cluster options include single H100 PCIe at $2.50/hr through to 1,024-GPU B200 InfiniBand clusters. The B200 SXM at $3.75/hr delivers roughly 3x the H200 throughput on large models per MLPerf reporting, which changes the self-hosting math further for teams at scale.

Frequently asked questions

The absolute cheapest paid LLM API is Groq Llama 3.1 8B at $0.05 per million input tokens and $0.08 per million output tokens. For a capable general-purpose model, DeepSeek V4 Flash on third-party providers like Together AI costs $0.14/$0.28 per million tokens at the old flat rate. DeepSeek direct switched to peak/off-peak billing on August 16 - off-peak is $0.22/$0.66. packet.ai Token Factory serves Llama 3.1 8B at $0.06/M and Llama 3.3 70B at $0.59/M with no cold starts.
On August 16, 2026, DeepSeek activated peak/off-peak billing on all V4 models, replacing the flat rate structure. V4 Flash moved from $0.14/$0.28 flat to $0.22/$0.66 off-peak and $0.44/$1.32 during peak windows (01:00-04:00 and 06:00-10:00 UTC). The change reflects demand-based pricing tied to the V4 official release. Third-party providers hosting the V4 Flash 0731 checkpoint - including Together AI - still offer the old flat rate. See the full DeepSeek V4 Flash pricing guide for routing strategies.
At 1 million tokens per day (input + output combined), daily cost at August 2026 rates: Groq Llama 3.1 8B ~$0.07/day, packet.ai Token Factory Llama 3.3 70B ~$0.59/day, DeepSeek V4 Flash off-peak ~$0.44/day (blended input/output), GPT-5.6 Luna ~$0.70/day, Claude Haiku 4.5 ~$3/day, Claude Opus 5 ~$15/day. Output-heavy workloads cost significantly more due to the 3-6x output multiplier across all providers.
Yes, significantly. DeepSeek V4 Flash off-peak costs $0.22/$0.66 per million tokens. OpenAI's cheapest current-generation model, GPT-5.6 Luna, costs $0.20/$1.20. On input tokens they are comparable; DeepSeek wins on output. Against GPT-5.6 Sol ($5/$30) or Claude Opus 5 ($5/$25), DeepSeek V4 Flash is 95-97% cheaper on input. The quality gap also narrowed significantly - DeepSeek V4 Flash scores 81% on SWE-bench Verified, within range of most flagship models for coding tasks.
For Llama 3.1 8B, Groq is cheapest at $0.05/$0.08 per million tokens. packet.ai Token Factory is $0.06/M with no cold starts and no free-tier rate limit constraints. For Llama 3.3 70B, packet.ai Token Factory at $0.59/M is 43% below Together AI's $1.04/M for the same model. For Llama 4 Maverick, Together AI and Fireworks serve it at approximately $0.15/$0.60. Meta does not sell a first-party Llama API - all rates come from third-party hosts. See Token Factory for current model availability.
For Llama 3.3 70B on a single H100 PCIe at packet.ai ($2.50/hr), the crossover against packet.ai Token Factory ($0.59/M) is approximately 12 million output tokens per month at typical utilisation. For the full 671B DeepSeek V3 open-weight model requiring an 8-GPU H200 node (~$1,800/month), the crossover against DeepSeek V4 Flash off-peak pricing is 2-3 billion tokens per month. Below these thresholds, managed APIs are cheaper and involve no infrastructure management. See the LLM inference cost breakdown for the full crossover math.
Three levers cut a real bill without changing the model. First, enable the Batch API where available (OpenAI, Anthropic, Google) - 50% off for async workloads with up to 24-hour turnaround. Second, enable prompt caching for any system prompt or document repeated across requests: cache reads cost 0.1x the base input rate at Anthropic and a similar fraction at DeepSeek. Third, set explicit max_tokens limits on responses - without them, models generate longer responses than needed, billing output tokens you did not want. Stack all three and costs can drop 60-75% on the same model at the same quality.

Last reviewed: August 24, 2026. Pricing verified against DeepSeek API documentation (August 21, 2026 via morphllm.com), IntuitionLabs LLM pricing comparison (August 19, 2026), Spheron Network LLM pricing comparison, and official provider pricing pages. LLM API pricing changes frequently - verify on provider pricing pages before committing to a budget. For managed GPU inference on open models, see packet.ai Token Factory. For self-hosted inference infrastructure, browse packet.ai cluster options.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog