Meta Llama 3.1

Llama 3.1 405B Instruct

Meta's 405B instruction-tuned Llama 3.1 model. The published throughput rows use NVIDIA's FP8 checkpoint across eight H100 SXM 80 GB GPUs.

Profile checked · Official model page ↗

Parameters
405B
Context
131072
Min. VRAM
—

Hardware evidence

Benchmarked GPU profiles

Benchmark GPUArchitectureVRAMRental observed from
NVIDIA H100 SXM 80 GBHopper80 GB$2.15/hr

Measured performance

Inference benchmarks

Model GPU setup Engine Quant. Workload Output tok/s Evidence
Llama 3.1 405B Instruct H100 SXM 80 GB8 GPUs TensorRT-LLM PyTorch FP8 128 → 2048input → output 4517.4 Source ↗Published Sep 15, 2025
Llama 3.1 405B Instruct H100 SXM 80 GB8 GPUs TensorRT-LLM PyTorch FP8 1000 → 1000input → output 2955.5 Source ↗Published Sep 15, 2025
Llama 3.1 405B Instruct H100 SXM 80 GB8 GPUs TensorRT-LLM PyTorch FP8 2048 → 128input → output 433.5 Source ↗Published Sep 15, 2025

Output throughput is aggregate maximum-load performance, not single-user generation speed. Workloads show input → output tokens. Compare rows only when the model, GPU count, quantization, engine, and workload match.