🚀 B200 starting at $3.75/hr. The best price you'll find. DC in US West → (Access it from button on top after login).

Get Your B200 →
Start Building
packet.ai/GPUs/NVIDIA L40S
In stock · Provisions in ~5 min
NVIDIA Ada Lovelace48 GB GDDR6PCIe Gen4

NVIDIA L40S GPU

Ada Lovelace for inference.

The NVIDIA L40S is the Ada Lovelace inference GPU, 48 GB of GDDR6 memory with Ada's 4th-gen Tensor Cores, delivering 733 TFLOPS of FP16 compute at a fraction of H100 cost. Ideal for cost-efficient LLM inference and generative AI. Available on packet.ai from $0.92/GPU-hour.

from $0.92/GPU-hr· Dedicated $0.92/hr · Monthly from $604/mo
≈ 76% below H100 cost
No contractsHourly billingSSH in <5 min
NVIDIA L40S GPU
48GB
GDDR6 memory
864GB/s
Memory bandwidth
733TFLOPS
FP16 Tensor compute
91.6TFLOPS
FP32 compute
Specifications

NVIDIA L40S specifications.

SpecificationValueGreat for
GPU architecture
NVIDIA Ada LovelaceAD102 · 76B transistors
Same Ada die as RTX 4090, tuned for 24/7 data-center operation.
GPU memory
48 GBGDDR6 with ECC
7B–13B LLMs natively; 30B+ at 4-bit.
Memory bandwidth
864 GB/s
Sufficient for small-batch inference.
FP16 Tensor compute
733 TFLOPSwith sparsity: 1,457
Higher FP16 throughput than A100 at lower cost.
FP32 compute
91.6 TFLOPS
Strong general-purpose compute.
Tensor Core gen
4th generationFP8 support
FP8 inference precision for LLM serving.
AV1 / RT Cores
YesAV1 encode + RT Gen 3
Hardware video encoding for AI video.
Form factor
PCIe Gen4Single-slot
No SXM motherboard required.
Availability
On-demandUS & EU regions
SSH-ready in under 5 minutes.
Architecture

Ada Lovelace for data-center inference.

The L40S brings Ada's 4th-gen Tensor Cores, FP8 precision, and AV1 hardware encode to inference-focused server workloads, at GDDR6 price points.

4th-gen Tensor Cores + FP8

Ada Tensor Cores with FP8 deliver up to 733 TFLOPS of FP16 compute. Strong throughput at a fraction of H100 cost.

48 GB GDDR6 at 864 GB/s

48 GB fits 7B to 13B models natively; 30B+ at 4-bit.

AV1 encode / RT Cores

Hardware AV1 encoding and 3rd-gen RT Cores make the L40S the best GPU for AI video generation.

PCIe Gen4, server-ready

Single-slot PCIe drops into any server without SXM motherboards.

Use Cases

What the L40S is built for.

Cost-efficient LLM inference

7B–13B fit natively; 30B+ at 4-bit. High FP16 throughput at under half H100 cost.

  • $0.92/hr dedicated
  • 733 TFLOPS FP16
  • vLLM / TGI optimised

Image & video generation

AV1 hardware encode and Ada RT Cores make L40S ideal for SDXL, FLUX, and video generation.

  • AV1 hardware encode
  • 3rd-gen RT Cores
  • SDXL / FLUX native

Fine-tuning on a budget

LoRA and QLoRA fine-tuning of 7B–13B models at the lowest cost per experiment.

  • $0.92/hr dedicated
  • 48 GB for fine-tuning
  • Hourly billing
Detailed Pricing Options

View all pricing tiers and configurations for L40S

ConfigurationOn-DemandMonthly3 Months6 MonthsAnnually
DedicatedDedicated only
2× NVIDIA L40SDedicated only

$1.84/hr

$1,142/mo

$1.56/hr eff.

$3,426/3mo

$6,852/6mo

$13,704/yr

4× NVIDIA L40SDedicated only

$3.68/hr

$2,283/mo

$3.13/hr eff.

$6,849/3mo

$13,698/6mo

$27,396/yr

Deploy hourly →Subscribe to Monthly →
FAQ

NVIDIA L40S, answered.

What is the NVIDIA L40S?

The L40S is NVIDIA's Ada Lovelace inference GPU, 48 GB GDDR6, 733 TFLOPS FP16, AV1 encode, and FP8 Tensor Cores.

How much does an L40S cost?

L40S starts at $0.92/GPU-hour dedicated. Monthly from $604/mo.

L40S vs H100?

H100 has HBM3 memory (3.35 TB/s vs 864 GB/s) and higher sustained throughput. L40S is cheaper and better for small-batch inference and image generation.

What models fit in L40S?

7B and 13B at FP16 fit easily. 30B–70B at 4-bit. For full FP16 70B, use A100 or H100.

Is L40S good for image generation?

Yes, one of the best value GPUs for Stable Diffusion XL, FLUX, and video generation.

Deploy now

Run the L40S. Ada inference from $0.92/hr.

Production-grade inference at $0.92/hr dedicated or $604/mo flat.

On-demand · hourly billing · US & EU regions

NVIDIA L40Sfrom $0.92/GPU-hr
Deploy L40S →