Meta Llama 3.1
Llama 3.1 8B Instruct
Meta's 8B instruction-tuned Llama 3.1 model. The published throughput rows use NVIDIA's FP8 checkpoint on one H100 SXM 80 GB GPU.
Profile checked · Official model page ↗
- Parameters
- 8B
- Context
- 131072
- Min. VRAM
- —
Hardware evidence
Benchmarked GPU profiles
| Benchmark GPU | Architecture | VRAM | Rental observed from |
|---|---|---|---|
| NVIDIA H100 SXM 80 GB | Hopper | 80 GB | $2.15/hr |
Measured performance
Inference benchmarks
| Model | GPU setup | Engine | Quant. | Workload | Output tok/s | Evidence |
|---|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | H100 SXM 80 GB1 GPU | TensorRT-LLM PyTorch | FP8 | 128 → 2048input → output | 21413.2 | Source ↗Published Sep 15, 2025 |
| Llama 3.1 8B Instruct | H100 SXM 80 GB1 GPU | TensorRT-LLM PyTorch | FP8 | 1000 → 1000input → output | 14991.6 | Source ↗Published Sep 15, 2025 |
| Llama 3.1 8B Instruct | H100 SXM 80 GB1 GPU | TensorRT-LLM PyTorch | FP8 | 2048 → 128input → output | 3275.6 | Source ↗Published Sep 15, 2025 |
Output throughput is aggregate maximum-load performance, not single-user generation speed. Workloads show input → output tokens. Compare rows only when the model, GPU count, quantization, engine, and workload match.