Skip to main content
Deep Learning GPU Benchmarks 2026: RTX 5090 vs A100 vs H100

Image: NVIDIA Tesla A100 by NVIDIA, CC BY-SA 4.0, via Wikimedia Commons

ยท Last updated on

Deep Learning GPU Benchmarks 2026: RTX 5090 vs A100 vs H100


About prices:prices on this page are US street prices in USD, last checked October 2026. They are for reference only. Local prices differ by region and usually include VAT or sales tax, and availability changes quickly, so check the retailer before you buy.

Buying guides tell you which GPU to buy. This post tells you why, with training throughput numbers, inference speeds, and VRAM requirements across every major card in 2026. Whether you are choosing between an RTX 5090 and an H100 rental or deciding if an RTX 5060 Ti is enough for your use case, the numbers below will guide that decision.

Benchmark Methodology

The RTX 5090 is set as the 100-point baseline. Read these as estimates, not lab measurements. We did not run these tests ourselves. Inference figures are anchored to published single-stream Ollama measurements where they exist (for example Database Martโ€™s RTX 5090 tests: Llama 3.1 8B at about 150 tokens/sec, and a 32B model at about 45 tokens/sec on both the RTX 5090 and the H100). Training figures assume PyTorch 2.x with mixed precision (BF16/FP16) and are scaled from published results and each cardโ€™s specs, so treat them as a ranking with rough ratios. We found no published head-to-head training benchmark covering all of these cards: the H100 and A100 training scores rest on reports that an H100 runs about 1.2 to 1.6 times an RTX 5090 on large-batch work, and that the RTX 5090 beats the A100 in half- and single-precision training. Run your own workload before a large purchase.

One result surprises people: a single RTX 5090 is as fast as an H100 for single-user inference on models that fit in 32 GB, and faster than an A100. Datacenter cards win on VRAM (80 GB), multi-GPU scaling, and batched serving, not on single-stream speed.

Key metrics tracked:

  • Training throughput: tokens/second or images/second at a fixed batch size
  • Inference speed: tokens/second for autoregressive generation
  • Memory bandwidth: the true bottleneck for most modern workloads
  • VRAM ceiling: what models fit without CPU offloading

Why relative numbers: Absolute benchmark figures shift with driver updates, framework versions, and batch sizes. Relative performance between cards is far more stable and more useful for buying decisions.

GPU Performance Ladder

$2.49/hr)$1.99/hr)$1,650-1,850$1,100-1,250$750-850$400-500
GPUVRAMBandwidthTraining Score (est.)Inference ScorePrice Tier
H100 SXM5 80GB80 GB HBM33,350 GB/s~140100 Cloud only (
A100 SXM4 80GB80 GB HBM2e2,000 GB/s~8077 Cloud only (
RTX 5090 32GB32 GB GDDR71,792 GB/s100100~$5,000
RTX 5080 16GB16 GB GDDR7960 GB/s5588
RTX 5070 Ti 16GB16 GB GDDR7896 GB/s4275
RTX 5060 Ti 16GB16 GB GDDR7448 GB/s2540
RTX 3060 12GB12 GB GDDR6360 GB/s1425

Reading the table: Training score is most relevant for model fine-tuning and pretraining. Inference score matters for local LLM serving. Note that the RTX 5070 Ti scores higher on inference than training relative to the 5080, because inference is more bandwidth-bound and both cards share similar bandwidth despite the 5080 having more compute cores.

LLM Training Benchmarks

Training large language models is the most demanding deep learning workload. It requires both high VRAM capacity and high memory bandwidth, since gradient updates touch every parameter on every step.

Transformer Training Throughput (BF16, PyTorch 2.x, estimates)

GPUGPT-2 (117M) tok/sLlama 7B (LoRA) tok/sLlama 13B (LoRA) tok/sNotes
H100 SXM5~38,000~2,600~1,150More VRAM and NVLink scaling, not 3x the speed
A100 SXM4~21,500~1,480~650Slower than a 5090 per card; wins on 80 GB and ECC
RTX 5090~27,000~1,850~820Best consumer card
RTX 5080~15,000~1,020~450VRAM limit at 13B+
RTX 5070 Ti~11,500~780~340Good value for 7B work
RTX 5060 Ti~6,700~460~200Learning/prototyping
RTX 3060 12GB~3,800~260Does not fitBudget starting point

VRAM Requirements by Model Size

This is the hard constraint that determines whether a GPU can even run your workload.

Model SizeFull Precision (FP16)QLoRA Fine-tuneMinimum GPU
7B parameters~14 GB~6 GBRTX 3060 12GB (QLoRA)
13B parameters~26 GB~10 GBRTX 5060 Ti 16GB (QLoRA)
30B parameters~60 GB~20 GBRTX 5090 32GB (QLoRA), A100 (FP16)
70B parameters~140 GB~48 GB2x A100 (FP16), H100 (QLoRA)
405B parameters~810 GB~270 GBMulti-node H100 cluster

QLoRA changes the game: Techniques like QLoRA (4-bit quantization with low-rank adapters) reduce VRAM requirements dramatically. You can fine-tune a 13B model on an RTX 5060 Ti 16GB that would otherwise need 26 GB. For most practical fine-tuning work, VRAM requirements in the QLoRA column apply.

Local LLM Inference Benchmarks

Inference is bandwidth-bound. The GPU that moves weights from VRAM to the compute units fastest wins. Single-stream speed is only part of it: the H100โ€™s real advantage is 80 GB of VRAM and batched serving for many users at once. For one user on a model that fits in 32 GB, an RTX 5090 keeps pace with it.

GPULlama 8B Q4 (tok/s)Llama 70B Q4 (tok/s)Qwen 32B Q4 (tok/s)Notes
H100 SXM5 80GB~150+~24~45All models fit fully in VRAM
A100 80GB~110~18-20~35All models fit fully in VRAM
RTX 5090 32GB~150Does not fit (about 43 GB); a few tok/s with offload~4532B Q4 fits (~20 GB)
RTX 5080 16GB~130~2-4 (heavy offload)~15-20 (partial offload)Limited for 30B+ models
RTX 5070 Ti 16GB~105-125~2-4 (heavy offload)~12-18 (partial offload)Close to the 5080 on small models
RTX 5060 Ti 16GB~42-75~1-3 (heavy offload)~8-12 (partial offload)Usable for 7B-13B models
RTX 3060 12GB~35-45~1-2 (heavy offload)~4-7 (heavy offload)Budget entry point

Offloading penalty: When a model does not fit in VRAM, layers are offloaded to system RAM and loaded on demand over the PCIe bus. PCIe 5.0 maxes at around 64 GB/s vs GDDR7โ€™s 1,792 GB/s on the RTX 5090. Models that partially offload to RAM run at a fraction of full-VRAM speed. If you want to run 70B models locally at usable speeds, you need a professional card or a dual-GPU setup.

Image Generation Benchmarks

Image generation with FLUX and Stable Diffusion is a mix of compute and bandwidth. Higher VRAM enables larger batch sizes and higher resolution without tiling. The RTX 5090 rows come from a published community test (about 10 seconds per 20-step FLUX Dev image, and 10.1 it/s on SDXL). The other cards are estimates scaled from it, and published results for the same card differ by 2x or more depending on precision and software.

GPUFLUX Dev FP8 (1024px, it/s)SDXL FP16 (1024px, it/s)FLUX Dev FP16 (1024px)
RTX 5090 32GB~2.2 it/s (measured)~10 it/s (measured)~2.1 it/s (fits fully)
RTX 5080 16GB~1.5 it/s (est.)~5-7 it/sRequires tiling or CPU offload
RTX 5070 Ti 16GB~1.2 it/s (est.)~5-6 it/s (est.)Requires tiling or CPU offload
RTX 5060 Ti 16GB~0.7 it/s (est.)~3.5 it/s (est.)Requires tiling or CPU offload
RTX 3060 12GB~0.35 it/s (est.)~2 it/s (est.)Heavy tiling required

For image generation, the RTX 5080 hits a practical sweet spot. FLUX Dev FP8 runs well within its 16 GB, and SDXL is fast. The 5090โ€™s advantage for image gen is running FLUX Dev FP16 natively without memory tricks, which benefits professional workflows generating hundreds of images.

VRAM Requirements by Workload

12 GB VRAM (RTX 3060)

Good for: fine-tuning 7B models with QLoRA, running 7B models at full speed, SDXL image generation, computer vision training on standard datasets.

Not good for: anything 13B or larger without heavy quantization, FLUX Dev in full precision, multi-model inference pipelines.

16 GB VRAM (RTX 5060 Ti, 5070 Ti, 5080)

Good for: fine-tuning 13B models with QLoRA, running 13B models comfortably, FLUX Dev FP8, most computer vision and NLP research workloads.

Not good for: 30B+ models without offloading, FLUX Dev FP16, full-precision 13B fine-tuning.

32 GB VRAM (RTX 5090)

Good for: fine-tuning up to 30B models with QLoRA, running Qwen 32B Q4 fully in VRAM, FLUX Dev FP16 natively, multi-model inference setups, serious research workflows.

Not good for: 70B+ models at full precision, anything requiring NVLink between multiple GPUs for VRAM pooling.

80 GB VRAM (A100, H100)

Good for: pretraining and fine-tuning 70B models with full precision, production inference at scale, multi-tenant GPU serving. This is the professional tier where the hardware is rented, not owned.

Pair with NVLink for 160 GB combined VRAM in dual configurations.

Price/Performance Analysis

Raw performance is only half the story. For most users, price per token/second or price per training step matters more than absolute throughput.

Prices rechecked October 9, 2026. These are US street prices, not MSRPs. A GDDR7 memory shortage has pushed every RTX 50 card well above its launch price (the RTX 5090 launched at $1,999 and now sells for around $5,000), so the value ranking looks very different from early 2026. Prices move weekly. See our best GPU for deep learning guide for current picks by budget.

GPUApprox. PriceTraining ScorePerf per $1,000Verdict
RTX 3060 12GB~$4501431Cheapest way in, no longer a bargain
RTX 5060 Ti 16GB~$8002531Cheapest new 16 GB card
RTX 5070 Ti 16GB~$1,1504237Best new-card price/performance
RTX 5080 16GB~$1,7505531Faster, but worse value than the 5070 Ti
RTX 5090 32GB~$5,00010020Worst ratio by far; you pay for 32 GB

The VRAM premium: The RTX 5090 has the lowest performance-per-dollar of any consumer card above, but it is the only consumer option that handles 30B models and FLUX FP16. You are paying a premium not just for speed but for VRAM capacity. If your workload fits in 16 GB, the RTX 5070 Ti gives you almost twice the performance per dollar at under a quarter of the price.

Recommendations by Use Case

Learning and Coursework

RTX 3060 12GB or RTX 5060 Ti 16GB. The 3060 handles every PyTorch tutorial and Hugging Face course without issue. The 5060 Ti 16GB adds headroom for 13B models if you want to experiment with local LLMs too.

Research and 7B-13B Fine-tuning

RTX 5070 Ti 16GB or RTX 5080 16GB. Both give you comfortable 16 GB for QLoRA fine-tuning up to 13B, fast SDXL generation, and solid PyTorch training throughput. The 5070 Ti is the better price/performance pick; the 5080 gives you a bit more compute headroom.

Serious Local LLM and 30B Fine-tuning

RTX 5090 32GB. The 32 GB VRAM is the key advantage, not the compute. If you regularly work with 32B-class models at Q4 or FLUX Dev FP16, the 5090 is the only consumer card that handles these without painful CPU offloading. A 70B model at Q4 (about 43 GB) still does not fit; that takes two 24 GB cards or an 80 GB rental.

Enterprise and 70B+ Training

Cloud A100 or H100 via RunPod, Lambda Labs, or Vast.ai. No consumer card can handle full-precision 70B training. Rent cloud GPUs for training runs, use a local RTX 5090 for iteration and inference.

Frequently Asked Questions

Is the H100 worth buying outright for a home lab?

Almost certainly not. A used H100 SXM5 runs $25,000 or more and requires a server platform with NVLink. For the same money you could rent H100 cloud time for years. The H100 makes sense for companies running continuous multi-month training jobs, not for individual researchers.

Does the RTX 5090 compete with the A100?

Yes, and on speed it wins. In published single-stream Ollama tests on a 32B model, the RTX 5090 ran at about 45 tokens per second, level with an H100 and ahead of an A100 at about 35. The A100โ€™s advantages are 80 GB of VRAM, ECC memory, and NVLink for multi-GPU training, so for anything that needs more than 32 GB it wins outright on capacity.

Should I buy two RTX 5080s instead of one RTX 5090?

Two RTX 5080s give you 32 GB total VRAM (16 GB each, not pooled unless using NVLink which the 5080 does not support), roughly double the compute, but VRAM is not shared transparently. For training with model parallelism, this can work. For inference, most tools require the model to fit on a single card. One RTX 5090 is simpler and gives you 32 GB on a single card, with no multi-GPU setup to manage. No RTX 50 card supports NVLink. See our CPU guide for platform requirements if you do go multi-GPU.

How do AMD GPUs perform for deep learning in 2026?

AMDโ€™s ROCm stack has improved significantly. For inference with llama.cpp, AMD RX 9070 XT (16 GB) is competitive with NVIDIA equivalents at a lower price. For training with PyTorch, ROCm support is solid for standard workloads but custom CUDA kernels and some libraries still require NVIDIA hardware. For pure learning and standard training tasks, AMD is a viable option in 2026 if the price is right.

What about the RTX 5090 vs RTX 4090 for training?

The 5090 wins by 45-75% on training tasks thanks to 78% more memory bandwidth and 33% more VRAM. The 4090 remains competitive for 7B-13B workloads where 24 GB is sufficient, though at October 2026 used prices of about $2,600 to $3,000 it is no longer a bargain. For a full comparison see our RTX 5090 vs 4090 deep dive.

Ready to Pick Your GPU?

Need help picking the right system? See our AI Workstation Build Guide for GPU, CPU, RAM, and storage recommendations together.