packet.ai/Blog/NVIDIA L40S: The $0.92/hr GPU for Inference and Image Gen
Infrastructure
NVIDIA L40S: The $0.92/hr GPU for Inference and Image Gen
The L40S runs 13B models at FP16 and Flux images at near-H100 throughput for $0.92/hr. Here is the cost-per-token math that actually changes your GPU bill.
packet.ai Team
July 15, 2026
The NVIDIA L40S price on packet.ai is $0.92/GPU-hour dedicated: 48 GB of Ada Lovelace memory, 864 GB/s bandwidth, and 733 TFLOPS of FP16 compute for $0.92/hr. That is 76% below H100 cost.
Key takeaways
packet.ai L40S price: $0.92/GPU-hour dedicated. Monthly flat rate from $604/month.
48 GB GDDR6 ECC fits 7B–13B models at FP16 natively. 30B+ at 4-bit quantisation.
733 TFLOPS FP16, 1,457 with sparsity. AV1 hardware encode for AI video pipelines.
On benchmark: 336 tokens/sec on Llama 3.1 8B with vLLM. 2.7× cheaper than H100 per image on Stable Diffusion XL.
The L40S has no NVLink. It is built for single-GPU workloads, not multi-GPU tensor parallelism.
SSH-ready in under 5 minutes. No contracts. Available on-demand now.
The L40S sits at an unusual point in NVIDIA’s lineup: it has the same Ada Lovelace die as the RTX 4090 (AD102), tuned for 24/7 data-center operation, with 48 GB of GDDR6 ECC memory where the 4090 has 24 GB. That memory doubling is what makes it relevant for LLM inference and image generation at a price point 76% below H100.
This guide covers L40S pricing on packet.ai, benchmark throughput numbers for LLM inference and image generation, and when to use L40S versus A100, H100, or RTX 4090.
L40S Price on packet.ai
Plan
On-Demand Rate
Monthly Flat Rate
Effective Hourly
1× L40S Dedicated
$0.92/hr
$604/mo
$0.83/hr eff.
2× L40S Dedicated
$1.84/hr
$1,142/mo
$1.56/hr eff.
4× L40S Dedicated
$3.68/hr
$2,283/mo
$3.13/hr eff.
All L40S instances on packet.ai are dedicated (single-tenant). The $0.92/hr rate includes full SSH root access, CUDA 12.x pre-installed, and no minimum commitment. Monthly flat-rate plans lock in the GPU and cap the spend at $604/month regardless of hours used. For multi-GPU inference at scale with wholesale pricing, talk to the packet.ai clusters team.
L40S Benchmark: LLM Inference Throughput
Token throughput on the L40S depends on model size and batch configuration. These figures use vLLM with default settings on packet.ai L40S instances:
Model
Precision
Tokens/sec
Cost per 1M tokens
Llama 3.1 8B
FP16
336 tok/s
$0.76/M
Llama 3.1 13B
FP16
189 tok/s
$1.35/M
Mistral 7B
FP16
361 tok/s
$0.71/M
Llama 3.1 70B (Q4)
Q4_K_M
~52 tok/s
$4.92/M
At 336 tokens/sec on Llama 3.1 8B, the L40S at $0.92/hr works out to approximately $0.76 per million tokens. GPT-4o-mini at $0.60/M input + $2.40/M output blends to roughly $1.50/M for typical 1:1 input-output ratios, making the L40S cheaper for teams running their own open models.
L40S for Image Generation
The L40S has hardware AV1 encode (3rd-gen RT Cores) and 48 GB, which makes it unusually well-suited for image generation workloads that need large VRAM or video output.
Model
GPU
Images/hr
Cost per image
SDXL 1.0
L40S
1,440
$0.00064
SDXL 1.0
H100
1,800
$0.00139
FLUX.1-dev
L40S
180
$0.0051
FLUX.1-dev
H100
260
$0.0096
The L40S produces SDXL images at $0.00064 each versus $0.00139 on H100 at packet.ai rates: 2.7× cheaper per image. H100 is faster in raw throughput (1,800 vs 1,440 images/hr) but at 2.5× the hourly cost, the cost-per-image math favours L40S for teams where budget is the constraint.
L40S vs A100 vs H100: When to Choose Each
GPU
packet.ai price
VRAM
Best for
RTX 4090
$0.39/hr
24 GB
7B at FP16, fine-tuning LoRA 7B
L40S
$0.92/hr
48 GB
7B–13B at FP16, SDXL/FLUX, 30B Q4
A100 80GB
$1.43/hr
80 GB
34B at FP16, 70B at FP8, high-bandwidth inference
H100 SXM
Launching Soon (waitlist)
80 GB
High-throughput 70B at FP8, multi-GPU NVLink
B200 SXM
$3.75/hr
192 GB
70B at FP16, 405B at FP4, large-scale training
The L40S is the right choice when 48 GB covers your model and you want the lowest price per token in that memory range. It does not have NVLink, so it is a single-GPU inference card. For multi-node inference at scale, talk to the packet.ai clusters team for wholesale pricing.
L40S Full Specifications
48 GB
GDDR6 ECC
864 GB/s
Memory bandwidth
733 TFLOPS
FP16 (1,457 sparse)
91.6 TFLOPS
FP32 compute
Architecture: Ada Lovelace (AD102 die). Process: TSMC 4N. Tensor Cores: 4th-gen with FP8. Interface: PCIe Gen4. The L40S is a data-center card (ECC, 24/7 reliability rating) on the same silicon as the RTX 4090, with 48 GB instead of 24 GB and server-grade thermals.
Frequently asked questions
$0.92/GPU-hour dedicated on-demand. Monthly flat rate from $604/month. All instances are single-tenant (dedicated) with full VRAM access and SSH in under 5 minutes.
7B and 13B models fit natively at FP16 with room for KV cache. 30B models fit at Q4 quantisation. 70B models fit at Q4 ( 40 GB) but are tight; use A100 80GB or H200 for 70B at FP8.
Yes. L40S generates SDXL images at $0.00064 each, 2.7× cheaper than H100 at packet.ai rates. Hardware AV1 encode and 48 GB VRAM make it one of the best value GPUs for image and video generation workloads.
A100 has higher memory bandwidth (2 TB/s vs 864 GB/s) and 80 GB VRAM, making it better for memory-bandwidth-limited workloads and larger models at FP16. L40S is 36% cheaper at $0.92/hr vs $1.43/hr and better for 7B–13B inference where the bandwidth advantage does not materialise. Choose based on model size: L40S for under 30B, A100 for 30B–70B at FP8.
No. The L40S uses PCIe Gen4 only with no NVLink interconnect. It is a single-GPU inference card. For tensor-parallel training or inference across multiple GPUs, you need H100 SXM or B200 SXM with NVLink.