Meta Llama 3.3

Llama 3.3 70B Instruct

Meta's 70B instruction-tuned Llama 3.3 model. The published throughput rows use NVIDIA's FP8 checkpoint across two H100 SXM 80 GB GPUs.

Profile checked · Official model page ↗

Parameters
70B
Context
131072
Min. VRAM
—

Hardware evidence

Benchmarked GPU profiles

Benchmark GPUArchitectureVRAMRental observed from
NVIDIA H100 SXM 80 GBHopper80 GB$2.15/hr

Measured performance

Inference benchmarks

Model GPU setup Engine Quant. Workload Output tok/s Evidence
Llama 3.3 70B Instruct H100 SXM 80 GB2 GPUs TensorRT-LLM PyTorch FP8 128 → 2048input → output 5892.9 Source ↗Published Sep 15, 2025
Llama 3.3 70B Instruct H100 SXM 80 GB2 GPUs TensorRT-LLM PyTorch FP8 1000 → 1000input → output 4181.1 Source ↗Published Sep 15, 2025
Llama 3.3 70B Instruct H100 SXM 80 GB2 GPUs TensorRT-LLM PyTorch FP8 2048 → 128input → output 723.4 Source ↗Published Sep 15, 2025

Output throughput is aggregate maximum-load performance, not single-user generation speed. Workloads show input → output tokens. Compare rows only when the model, GPU count, quantization, engine, and workload match.