# packet.ai > On-demand GPU cloud for AI and ML workloads. NVIDIA B200, H200, A100, > and RTX series GPUs from $0.39/hr. No contracts, deploy in under 5 minutes. > Built by hosted.ai. packet.ai provides infrastructure-grade GPU compute at 50%+ below hyperscaler pricing, powered by intelligent GPU scheduling. European-founded, US and EU datacenter regions, no reserved-capacity lock-in. ## Key pages - [Homepage](https://packet.ai/): On-demand GPU cloud overview and pricing entry - [Pricing](https://packet.ai/pricing): Full GPU pricing across Dynamic, Dedicated, and Clusters tiers - [B200 GPU](https://packet.ai/gpu/b200): NVIDIA B200 — 192GB HBM3e, from $3.75/hr Dynamic, $5.90/hr Dedicated - [H200 GPU](https://packet.ai/gpu/h200): NVIDIA H200 — 141GB HBM3e, from $2.49/hr Dynamic - [H100 GPU](https://packet.ai/gpu/h100): NVIDIA H100 SXM — 80GB HBM3, from $2.50/hr - [A100 GPU](https://packet.ai/gpu/a100): NVIDIA A100 80GB, from $1.43/hr Dedicated - [L40S GPU](https://packet.ai/gpu/l40s): NVIDIA L40S 48GB, from $0.92/hr Dedicated - [RTX 6000 Pro](https://packet.ai/gpu/rtx-6000): NVIDIA RTX PRO 6000 Blackwell 96GB, from $0.66/hr Dynamic - [RTX 4090](https://packet.ai/gpu/rtx-4090): NVIDIA RTX 4090 24GB, from $0.39/hr - [RTX 5090](https://packet.ai/gpu/rtx-5090): NVIDIA RTX 5090 32GB, from $0.59/hr - [Dedicated GPU Cloud](https://packet.ai/dedicated-gpu-cloud): Single-tenant GPU pods with 99.99% SLA - [Dynamic GPU Cloud](https://packet.ai/dynamic-gpu-cloud): Shared GPU pods, same peak performance, lower cost - [Bare Metal GPU](https://packet.ai/bare-metal-gpu-servers): Dedicated single-node bare metal GPU servers - [GPU Clusters](https://packet.ai/gpu-cluster): Multi-node GPU clusters with InfiniBand fabric - [Token Factory](https://packet.ai/token-factory): OpenAI-compatible LLM inference API - [Pixel Factory](https://packet.ai/pxl): Image and video generation API ## Comparisons - [packet.ai vs AWS](https://packet.ai/vs/aws-gpu): GPU pricing comparison — packet.ai vs AWS - [packet.ai vs CoreWeave](https://packet.ai/vs/coreweave): packet.ai vs CoreWeave - [packet.ai vs RunPod](https://packet.ai/vs/runpod): packet.ai vs RunPod - [packet.ai vs Vast.ai](https://packet.ai/vs/vast-ai): packet.ai vs Vast.ai - [packet.ai vs Lambda Labs](https://packet.ai/vs/lambda-labs): packet.ai vs Lambda Labs - [packet.ai vs Hyperstack](https://packet.ai/vs/hyperstack): packet.ai vs Hyperstack ## Blog - [What Is NVfp4? Blackwell's 4-Bit Precision Format Explained](https://packet.ai/blog/nvfp4): NVfp4 cuts model memory 3.5x vs FP16. B200 runs at 9,000 TFLOPS FP4 - [FP8 vs FP16 vs BF16: Precision Formats for AI Explained](https://packet.ai/blog/fp8-vs-fp16-vs-bf16): FP8 vs FP16 vs BF16 for AI training and inference - [NVLink vs PCIe for AI Workloads: When the Interconnect Is Your Bottleneck](https://packet.ai/blog/nvlink-vs-pcie): NVLink vs PCIe compared for LLM training and inference - [HBM3e vs HBM2e: Why Memory Bandwidth Decides LLM Throughput](https://packet.ai/blog/hbm3e-vs-hbm2e): HBM3e vs HBM2e and why memory bandwidth decides LLM inference throughput - [NVIDIA Blackwell Architecture Explained: What's New in the B200 and RTX 50 Series](https://packet.ai/blog/nvidia-blackwell-architecture): NVIDIA Blackwell dual-die GPU design with 208B transistors explained - [Flux Image Generation on Cloud GPUs: VRAM Requirements, Speed and Cost Per Image](https://packet.ai/blog/flux-image-generation-gpu-vram-requirements): Flux image generation VRAM requirements at FP8 on RTX 4090 - [RTX 5090 for AI in the Cloud: Pricing, Benchmarks vs RTX 4090](https://packet.ai/blog/rtx-5090-cloud-gpu-ai): RTX 5090 cloud GPU from $0.59/hr: 32GB GDDR7, vLLM benchmarks - [RTX 4090 for AI in the Cloud: Pricing, Performance and How It Compares to A100](https://packet.ai/blog/rtx-4090-for-ai-cloud): RTX 4090 for AI workloads: rental pricing from $0.39/hr, inference benchmarks - [L40S GPU Cloud: Choosing the Right GPU for Production Inference](https://packet.ai/blog/l40s-gpu-cloud-cheapest-inference): L40S at $0.92/GPU-hour: 48GB VRAM and FP8 Tensor Cores for inference - [vLLM Docker Deployment: Best GPU for Production LLM Inference](https://packet.ai/blog/vllm-docker-deployment-gpu-inference): vLLM Docker deployment guide for production LLM inference - [NVIDIA A100 vs H100 in 2026: Price, Performance and Which GPU Fits Your Workload](https://packet.ai/blog/nvidia-a100-vs-h100): A100 vs H100 specs, vLLM throughput benchmarks, QLoRA fine-tuning cost - [NVIDIA L40S: The $0.92/hr GPU for Inference and Image Gen](https://packet.ai/blog/nvidia-l40s-price-inference-image-generation): NVIDIA L40S pricing from $0.92/GPU-hr, 48GB Ada Lovelace - [Mixtral 8x22B GPU Requirements: The MoE Memory Math](https://packet.ai/blog/mixtral-8x22b-gpu-requirements): Mixtral 8x22B VRAM requirements for all 141B parameters - [Gemma 4 12B Deployment: Which Single GPU Actually Fits?](https://packet.ai/blog/gemma-4-12b-deployment): How to deploy Gemma 4 12B on a single GPU: real VRAM requirements - [AWQ vs GPTQ vs FP8: LLM Quantization Methods Compared for Serving](https://packet.ai/blog/awq-vs-gptq-vs-fp8): AWQ and GPTQ quantization compared to FP8 for LLM serving - [Speculative Decoding Explained: Draft Models and 2-3x Faster Inference](https://packet.ai/blog/speculative-decoding-explained): Speculative decoding speeds up LLM inference with draft models - [vLLM Prefix Caching Explained: Reduce Latency for Repeated Prompts](https://packet.ai/blog/vllm-prefix-caching): vLLM automatic prefix caching: hash-based mechanism and configuration - [SXM vs PCIe GPUs: Form Factors, Power Limits and Real Performance Gaps](https://packet.ai/blog/sxm-vs-pcie-gpus): SXM vs PCIe GPUs: H100 SXM 700W vs 350W PCIe, real performance gaps - [Continuous Batching Explained: How Modern Servers Keep GPUs Full](https://packet.ai/blog/continuous-batching-explained): Continuous batching and chunked prefill for LLM inference servers - [SGLang vs vLLM vs TensorRT-LLM: The 2026 Inference Engine Decision Guide](https://packet.ai/blog/sglang-vs-vllm-vs-tensorrt-llm): SGLang, vLLM, or TensorRT-LLM throughput, TTFT, and GPU cost comparison - [Qwen 3.6 VRAM Requirements: Both Models, One Consumer GPU](https://packet.ai/blog/qwen-3-6-vram-requirements): Qwen 3.6 27B and 35B-A3B VRAM requirements on consumer GPU - [What Is a GPU Cluster? Architecture, Networking and Scale Explained](https://packet.ai/blog/gpu-clusters-explained): GPU clusters from single-node to InfiniBand multi-node architecture - [LLM Inference Cost in 2026: API Pricing and Cost per Million Tokens Compared](https://packet.ai/blog/llm-inference-cost): LLM inference cost 2026: frontier APIs vs self-hosted Llama cost comparison - [Fine-Tuning LLMs on GPU Cloud: A100 vs H100, QLoRA Memory Math and Cost Per Run](https://packet.ai/blog/fine-tuning-llm-gpu-a100-vs-h100): QLoRA on A100 costs $34-$51 per run at $1.43/hr - [vLLM Tutorial: How to Deploy and Serve LLMs on GPU Cloud (2026)](https://packet.ai/blog/vllm-tutorial-deploy-serve-llms-gpu-cloud): Deploy vLLM on H100, H200, or B200 GPU cloud with PagedAttention - [Running Llama 3 and Llama 4 in the Cloud: GPU Sizing From 8B to 405B](https://packet.ai/blog/llama-gpu-sizing-8b-to-405b): GPU sizing for Llama 3 and Llama 4 from 8B to 405B with real VRAM figures - [Ollama in the Cloud: Why You'd Rent a GPU Instead of Running Locally](https://packet.ai/blog/ollama-cloud-gpu-vs-local): Ollama real VRAM needs and when cloud GPU beats local - [GPU Passthrough vs vGPU vs MIG: Which Mode Are You Actually Getting?](https://packet.ai/blog/gpu-passthrough-vs-vgpu-vs-mig): GPU passthrough, NVIDIA vGPU, and MIG isolation levels compared - [When Is a B200 Worth the Extra Cost? A B200 vs H200 ROI Framework](https://packet.ai/blog/b200-vs-h200-roi-framework): B200 vs H200 ROI math and real pricing comparison - [RTX PRO 6000 Blackwell: 96GB GPU for LLM Inference from $0.66/hr](https://packet.ai/blog/rtx-pro-6000-blackwell-gpu-cloud): RTX Pro 6000 rental from $0.66/hr with 96GB GDDR7 memory - [Why Your GPU Keeps Waiting: The CPU Bottleneck in Agentic AI Workloads](https://packet.ai/blog/cpu-bottleneck-agentic-ai): Agentic AI CPU bottleneck: tool calls leave GPUs idle 60%+ of the time - [H100 vs B200: Which GPU Is Right for Your AI Workload?](https://packet.ai/blog/h100-vs-b200-gpu-comparison): H100 vs B200 on specs, MLPerf benchmarks, and cloud pricing - [NVIDIA B200 GPU Cloud: Pricing, Specs & Where to Rent in 2026](https://packet.ai/blog/b200-gpu-cloud-pricing-specs): B200 GPU cloud pricing from $3.75/GPU-hr on packet.ai - [GLM-5.2 Self-Hosting Guide: GPU Sizing for the Top Open-Weight Coder](https://packet.ai/blog/glm-5-2-self-hosting-gpu-sizing-guide): GLM-5.2 needs 744GB at FP8, deployment guide for GPU cloud - [Not All GPU Hot Migration Is the Same](https://packet.ai/blog/gpu-hot-migration-shared-cloud-gpus): vGPU live migration vs CUDA workload migration compared - [Moonshot Kimi K3 GPU Requirements: Confirmed VRAM & Deployment Guide](https://packet.ai/blog/moonshot-kimi-k3-gpu-requirements): Kimi K3 1.56TB footprint and practical GPU cloud deployment guide - [DeepSeek V4 Pro & Flash GPU Requirements: What They Actually Need](https://packet.ai/blog/deepseek-v4-pro-flash-gpu-requirements): DeepSeek V4 Pro cluster requirements vs V4-Flash single B200 deployment - [Multi-Node Training on GPU Clusters: When One Server Stops Being Enough](https://packet.ai/blog/multi-node-gpu-cluster-training): NVLink within server vs InfiniBand between servers for multi-node training - [Video Generation in the Cloud: GPU Requirements for LTX, Hunyuan & Wan](https://packet.ai/blog/video-generation-gpu-requirements-ltx-hunyuan-wan): VRAM requirements for LTX-Video, HunyuanVideo, and Wan 2.1/2.2 - [Why We Raised $19M to Fix GPU Infrastructure](https://packet.ai/blog/why-we-raised-19m-to-fix-gpu-infrastructure): packet.ai $19M seed round led by Creandum to fix GPU infrastructure - [Langflow on a Cloud GPU: Size the Model, Not the Builder](https://packet.ai/blog/langflow-cloud-gpu-visual-llm-builder): Langflow GPU requirements and visual LLM builder on cloud - [GPU Cloud at 75% Less Than AWS: How packet.ai Prices Blackwell](https://packet.ai/blog/aws-gpu-pricing-vs-packet-ai-blackwell): AWS GPU pricing at $14.24/hr vs packet.ai Blackwell from $3.75/hr - [Bare Metal vs VM GPU Performance: What the Benchmarks Actually Show](https://packet.ai/blog/bare-metal-vs-vm-gpu-performance-benchmarks): Bare metal vs virtual machine GPU performance benchmarks - [Rent GPU for AI: VRAM Requirements Guide for Every Major Model (2026)](https://packet.ai/blog/rent-gpu-for-ai-vram-requirements-guide): GPU rental guide: L40S $0.92/hr, A100 $1.43/hr, B200 $3.75/hr with VRAM guide - [GPU Pricing Models Compared 2026: On-Demand vs Reserved vs Spot](https://packet.ai/blog/gpu-pricing-models-compared): On-demand, reserved, and spot GPU pricing models break-even comparison - [Why We Only Use Infrastructure Powered by hosted.ai](https://packet.ai/blog/nordic-datacenter-advantage): Every GPU on packet.ai runs on hosted.ai-powered infrastructure - [How Dynamic GPU Placement Enables Lower Prices](https://packet.ai/blog/dynamic-placement-explained): Dynamic GPU placement schedules workloads based on real-time VRAM and compute - [Welcome to Packet.ai](https://packet.ai/blog/introducing-packet-ai): packet.ai GPU cloud with H100, H200, B200, and RTX PRO 6000 at transparent pricing - [One-Click GPU Environments: VS Code, Jupyter, and More](https://packet.ai/blog/one-click-gpu-environments): One-click GPU dev environments: VS Code in browser, Jupyter Lab - [How Persistent Workspaces Actually Work (A Deep Dive)](https://packet.ai/blog/persistent-workspaces-deep-dive): Persistent Workspaces mount a PVC at /workspace for files and pip packages - [Token Factory: How We Built a 98% Cheaper OpenAI Alternative](https://packet.ai/blog/token-factory-deep-dive): Token Factory OpenAI-compatible inference at $0.10/M tokens - [Why We Built Token Factory](https://packet.ai/blog/why-we-built-token-factory): Token Factory is packet.ai's OpenAI-compatible inference API at $0.10/M tokens - [Packet.ai + SkyPilot: Run ML Workloads with One Command (Alpha)](https://packet.ai/blog/skypilot-integration-alpha): Native SkyPilot cloud provider for packet.ai ML workloads - [GPU Utilization: The Lie Your Dashboard Tells You](https://packet.ai/blog/gpu-utilization-the-lie-your-dashboard-tells): GPU_UTIL measures kernel activity, not actual compute utilization - [GPU Snapshots Demystified: What Actually Survives Pod Termination](https://packet.ai/blog/snapshots-and-storage-demystified): GPU pod snapshots: what survives pod termination - [30 Years Ago Today, a Computer Beat the World Chess Champion](https://packet.ai/blog/deep-blue-anniversary-gpus): Deep Blue 11.38 GFLOPS vs a single B200 GPU today - [8 Large Language Models on a Single NVIDIA Blackwell Server](https://packet.ai/blog/blackwell-rtx-pro-6000-inference): 8 open-source LLMs deployed on one 8x RTX PRO 6000 Blackwell server - [Text Generation WebUI (oobabooga) on a Cloud GPU: Full Setup Guide](https://packet.ai/blog/oobabooga-cloud-gpu-setup-guide): oobabooga TextGen on cloud GPU: VRAM requirements by model size - [The Noisy Neighbor Problem on GPU Cloud: p99 Latency, Throughput Variance](https://packet.ai/blog/noisy-neighbor-problem-gpu-cloud): Noisy neighbor problem on GPU cloud: p99 latency spikes explained - [ComfyUI in the Cloud: RTX 4090 vs L40S vs A100](https://packet.ai/blog/best-gpu-for-comfyui): Best GPU for ComfyUI: RTX 4090, L40S, or A100 80GB comparison - [Bare Metal GPU Servers vs Virtualized GPU Cloud: When Isolation Is Worth Paying For](https://packet.ai/blog/bare-metal-gpu-server-vs-virtualized-gpu-cloud): Bare metal GPU server vs virtualized GPU cloud: passthrough hits 96-100% of native ## Company - [About](https://packet.ai/about): Founded by hosted.ai — infrastructure veterans since 1996 - [Technology](https://packet.ai/technology): How intelligent GPU scheduling works - [Features](https://packet.ai/features): Platform capabilities - [Documentation](https://dash.packet.ai/docs): API reference, quick-start guides, SSH access