No items found.
Start Building
Infrastructure

NVIDIA L40S: The $0.92/hr GPU for Inference and Image Gen

The L40S runs 13B models at FP16 and Flux images at near-H100 throughput for $0.92/hr. Here is the cost-per-token math that actually changes your GPU bill.

Author photo
packet.ai Team
July 15, 2026

The NVIDIA L40S price on packet.ai is $0.92/GPU-hour dedicated: 48 GB of Ada Lovelace memory, 864 GB/s bandwidth, and 733 TFLOPS of FP16 compute for $0.92/hr. That is 76% below H100 cost.

Key takeaways

  • packet.ai L40S price: $0.92/GPU-hour dedicated. Monthly flat rate from $604/month.
  • 48 GB GDDR6 ECC fits 7B–13B models at FP16 natively. 30B+ at 4-bit quantisation.
  • 733 TFLOPS FP16, 1,457 with sparsity. AV1 hardware encode for AI video pipelines.
  • On benchmark: 336 tokens/sec on Llama 3.1 8B with vLLM. 2.7× cheaper than H100 per image on Stable Diffusion XL.
  • The L40S has no NVLink. It is built for single-GPU workloads, not multi-GPU tensor parallelism.
  • SSH-ready in under 5 minutes. No contracts. Available on-demand now.

The L40S sits at an unusual point in NVIDIA’s lineup: it has the same Ada Lovelace die as the RTX 4090 (AD102), tuned for 24/7 data-center operation, with 48 GB of GDDR6 ECC memory where the 4090 has 24 GB. That memory doubling is what makes it relevant for LLM inference and image generation at a price point 76% below H100.

This guide covers L40S pricing on packet.ai, benchmark throughput numbers for LLM inference and image generation, and when to use L40S versus A100, H100, or RTX 4090.

L40S Price on packet.ai

Plan On-Demand Rate Monthly Flat Rate Effective Hourly
1× L40S Dedicated $0.92/hr $604/mo $0.83/hr eff.
2× L40S Dedicated $1.84/hr $1,142/mo $1.56/hr eff.
4× L40S Dedicated $3.68/hr $2,283/mo $3.13/hr eff.

All L40S instances on packet.ai are dedicated (single-tenant). The $0.92/hr rate includes full SSH root access, CUDA 12.x pre-installed, and no minimum commitment. Monthly flat-rate plans lock in the GPU and cap the spend at $604/month regardless of hours used. For multi-GPU inference at scale with wholesale pricing, talk to the packet.ai clusters team.

L40S Benchmark: LLM Inference Throughput

Token throughput on the L40S depends on model size and batch configuration. These figures use vLLM with default settings on packet.ai L40S instances:

Model Precision Tokens/sec Cost per 1M tokens
Llama 3.1 8B FP16 336 tok/s $0.76/M
Llama 3.1 13B FP16 189 tok/s $1.35/M
Mistral 7B FP16 361 tok/s $0.71/M
Llama 3.1 70B (Q4) Q4_K_M ~52 tok/s $4.92/M

At 336 tokens/sec on Llama 3.1 8B, the L40S at $0.92/hr works out to approximately $0.76 per million tokens. GPT-4o-mini at $0.60/M input + $2.40/M output blends to roughly $1.50/M for typical 1:1 input-output ratios, making the L40S cheaper for teams running their own open models.

L40S for Image Generation

The L40S has hardware AV1 encode (3rd-gen RT Cores) and 48 GB, which makes it unusually well-suited for image generation workloads that need large VRAM or video output.

Model GPU Images/hr Cost per image
SDXL 1.0 L40S  1,440 $0.00064
SDXL 1.0 H100  1,800 $0.00139
FLUX.1-dev L40S  180 $0.0051
FLUX.1-dev H100  260 $0.0096

The L40S produces SDXL images at $0.00064 each versus $0.00139 on H100 at packet.ai rates: 2.7× cheaper per image. H100 is faster in raw throughput (1,800 vs 1,440 images/hr) but at 2.5× the hourly cost, the cost-per-image math favours L40S for teams where budget is the constraint.

L40S vs A100 vs H100: When to Choose Each

GPU packet.ai price VRAM Best for
RTX 4090$0.39/hr24 GB7B at FP16, fine-tuning LoRA 7B
L40S$0.92/hr48 GB7B–13B at FP16, SDXL/FLUX, 30B Q4
A100 80GB$1.43/hr80 GB34B at FP16, 70B at FP8, high-bandwidth inference
H100 SXMLaunching Soon (waitlist)80 GBHigh-throughput 70B at FP8, multi-GPU NVLink
B200 SXM$3.75/hr192 GB70B at FP16, 405B at FP4, large-scale training

The L40S is the right choice when 48 GB covers your model and you want the lowest price per token in that memory range. It does not have NVLink, so it is a single-GPU inference card. For multi-node inference at scale, talk to the packet.ai clusters team for wholesale pricing.

L40S Full Specifications

48 GB

GDDR6 ECC

864 GB/s

Memory bandwidth

733 TFLOPS

FP16 (1,457 sparse)

91.6 TFLOPS

FP32 compute

Architecture: Ada Lovelace (AD102 die). Process: TSMC 4N. Tensor Cores: 4th-gen with FP8. Interface: PCIe Gen4. The L40S is a data-center card (ECC, 24/7 reliability rating) on the same silicon as the RTX 4090, with 48 GB instead of 24 GB and server-grade thermals.

Frequently asked questions

$0.92/GPU-hour dedicated on-demand. Monthly flat rate from $604/month. All instances are single-tenant (dedicated) with full VRAM access and SSH in under 5 minutes.
7B and 13B models fit natively at FP16 with room for KV cache. 30B models fit at Q4 quantisation. 70B models fit at Q4 ( 40 GB) but are tight; use A100 80GB or H200 for 70B at FP8.
Yes. L40S generates SDXL images at  $0.00064 each, 2.7× cheaper than H100 at packet.ai rates. Hardware AV1 encode and 48 GB VRAM make it one of the best value GPUs for image and video generation workloads.
A100 has higher memory bandwidth (2 TB/s vs 864 GB/s) and 80 GB VRAM, making it better for memory-bandwidth-limited workloads and larger models at FP16. L40S is 36% cheaper at $0.92/hr vs $1.43/hr and better for 7B–13B inference where the bandwidth advantage does not materialise. Choose based on model size: L40S for under 30B, A100 for 30B–70B at FP8.
No. The L40S uses PCIe Gen4 only with no NVLink interconnect. It is a single-GPU inference card. For tensor-parallel training or inference across multiple GPUs, you need H100 SXM or B200 SXM with NVLink.

Last reviewed: July 2026. Deploy an L40S on packet.ai from $0.92/hr.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog