No items found.
Start Building
packet.ai/GPUs/NVIDIA RTX 4090
In stock · Provisions in ~5 min
NVIDIA Ada Lovelace24 GB GDDR6XPCIe Gen4

NVIDIA RTX 4090 GPU

The fastest consumer GPU for AI workloads.

The NVIDIA RTX 4090 brings 4th-gen Tensor Cores, FP8 support, and 24 GB GDDR6X to a PCIe card. The best value Ada GPU for AI inference and fine-tuning. Available from $0.39/GPU-hour.

from $0.39/GPU-hr· Dedicated $0.39/hr · Monthly $263/mo
Best value Ada GPU
No contractsHourly billingSSH in <5 min
NVIDIA RTX 4090 GPU
24GB
GDDR6X memory
1.01TB/s
Memory bandwidth
82.6TFLOPS
FP32 compute
1,321TOPS
AI compute (FP8)
Specifications

NVIDIA RTX 4090 specifications.

SpecificationValueGreat for
GPU architecture
NVIDIA Ada LovelaceAD102 · 76B transistors
Same die family as L40S, tuned for consumer AI workloads.
GPU memory
24 GBGDDR6X
7B at FP16 natively; 13B at Q4 on a single card.
Memory bandwidth
1.01 TB/s
Fast consumer memory for inference.
FP32 compute
82.6 TFLOPS
Strong general-purpose compute.
AI compute (FP8/INT8)
up to 1,321 AI TOPS
High-throughput low-precision inference.
Tensor Core gen
4th generation512 Tensor Cores
FP8 support for efficient inference.
CUDA cores
16,384
Massive parallel compute.
Host interface
PCIe Gen4x16, no NVLink
Fits any Gen4 server, no SXM board needed.
Availability
On-demandUS & EU regions
SSH-ready in under 5 minutes.
Architecture

Ada Lovelace, built for AI.

The RTX 4090 brings 4th-gen Tensor Cores, FP8 support, and 24 GB GDDR6X to a PCIe card.

4th-gen Tensor Cores

Ada Lovelace flagship consumer GPU with FP8 support and 1,321 AI TOPS.

24 GB GDDR6X

Enough memory for 7B at FP16 and 13B at Q4 quantisation on a single card.

PCIe Gen4

Drops into standard PCIe servers, no SXM motherboard required.

FP8 inference

4th-gen Tensor Cores with FP8 precision for efficient low-latency inference.

Use Cases

What the RTX 4090 is built for.

7B LLM inference

7B models fit natively at FP16. The most cost-effective GPU per token for small models.

  • $0.39/hr dedicated
  • 24 GB GDDR6X
  • vLLM / TGI ready

Fine-tuning & LoRA

LoRA and QLoRA fine-tuning of 7B models at the lowest hourly rate on packet.ai.

  • Lowest price in lineup
  • FP8 Tensor Cores
  • Hourly billing

Image generation

SDXL and FLUX run natively. Ada's RT Cores accelerate diffusion pipelines.

  • SDXL / FLUX native
  • AV1 hardware encode
  • $263/mo flat
Detailed Pricing Options

View all pricing tiers and configurations for RTX 4090

ConfigurationOn-DemandMonthly3 Months6 MonthsAnnually
DedicatedDedicated only
2× NVIDIA RTX 4090Dedicated only

$0.77/hr

$511/mo

$0.70/hr eff.

$1,533/3mo

$3,066/6mo

$6,132/yr

Deploy hourly →Subscribe to Monthly →
FAQ

NVIDIA RTX 4090, answered.

What is the NVIDIA RTX 4090?

Ada Lovelace flagship consumer GPU: 24 GB GDDR6X, 1.01 TB/s, 1,321 AI TOPS. Best value Ada GPU on packet.ai.

How much does RTX 4090 cost?

RTX 4090 starts at $0.39/GPU-hour dedicated, or $263/month flat.

Is RTX 4090 good for fine-tuning?

Yes. LoRA and QLoRA fine-tuning of 7B models fits comfortably in 24 GB.

What models fit in RTX 4090?

7B at FP16 natively; 13B at Q4 on a single card. For 30B+, use L40S or A100.

Does RTX 4090 support NVLink?

No. PCIe Gen4 only. For NVLink multi-GPU, use H100 SXM or B200.

Deploy now

Run the RTX 4090. 24 GB Ada from $0.39/hr.

The best value consumer GPU on packet.ai. $0.39/hr dedicated, or $263/mo flat.

On-demand · hourly billing · US & EU regions

NVIDIA RTX 4090from $0.39/GPU-hr
Deploy RTX 4090 →