The NVIDIA RTX 4090 brings 4th-gen Tensor Cores, FP8 support, and 24 GB GDDR6X to a PCIe card. The best value Ada GPU for AI inference and fine-tuning. Available from $0.39/GPU-hour.

The RTX 4090 brings 4th-gen Tensor Cores, FP8 support, and 24 GB GDDR6X to a PCIe card.
Ada Lovelace flagship consumer GPU with FP8 support and 1,321 AI TOPS.
Enough memory for 7B at FP16 and 13B at Q4 quantisation on a single card.
Drops into standard PCIe servers, no SXM motherboard required.
4th-gen Tensor Cores with FP8 precision for efficient low-latency inference.
7B models fit natively at FP16. The most cost-effective GPU per token for small models.
LoRA and QLoRA fine-tuning of 7B models at the lowest hourly rate on packet.ai.
SDXL and FLUX run natively. Ada's RT Cores accelerate diffusion pipelines.
| Configuration | On-Demand | Monthly | 3 Months | 6 Months | Annually |
|---|---|---|---|---|---|
| Dedicated | Dedicated only | ||||
| 1× NVIDIA RTX 4090Most Popular | $0.39/hr | $263/mo $0.36/hr eff. | $789/3mo | $1,578/6mo | $3,156/yr |
| 2× NVIDIA RTX 4090Dedicated only | $0.77/hr | $511/mo $0.70/hr eff. | $1,533/3mo | $3,066/6mo | $6,132/yr |
| Deploy hourly → | Subscribe to Monthly → |
Ada Lovelace flagship consumer GPU: 24 GB GDDR6X, 1.01 TB/s, 1,321 AI TOPS. Best value Ada GPU on packet.ai.
RTX 4090 starts at $0.39/GPU-hour dedicated, or $263/month flat.
Yes. LoRA and QLoRA fine-tuning of 7B models fits comfortably in 24 GB.
7B at FP16 natively; 13B at Q4 on a single card. For 30B+, use L40S or A100.
No. PCIe Gen4 only. For NVLink multi-GPU, use H100 SXM or B200.
The best value consumer GPU on packet.ai. $0.39/hr dedicated, or $263/mo flat.
On-demand · hourly billing · US & EU regions