Qwen3
Qwen3-32B
Qwen3 checkpoint with sourced BF16 weights and full-attention cache configuration. Memory estimates exclude unverified runtime overhead.
Profile checked · Official model page ↗
- Parameters
- 32.8B
- Context
- 40960
- Min. VRAM
- —
Workload memory estimate
- Checkpoint weights
- 61.02 GiB
- KV cache
- 2.00 GiB
- Weights + cache
- 63.02 GiB
BF16 weights / BF16 cache · 8192 cached tokens per sequence · 1 resident sequence.
Pinned checkpoint ↗ · 9216db5781bf21249d130ec9da846c4624c16137 · huggingface.co: config.json ↗ · huggingface.co: model.safetensors.index.json ↗
Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility or pooled multi-GPU memory. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.
Hardware evidence
Benchmarked GPU profiles
| Benchmark GPU | Architecture | VRAM | Rental observed from |
|---|---|---|---|
| No benchmarked GPU profiles are available yet. | |||
Measured performance
Inference benchmarks
| Model | GPU setup | Engine | Quant. | Workload | Reported tok/s | Evidence | Max-load capacity scenarioUSD / million output tokens |
|---|---|---|---|---|---|---|---|
| No benchmark records have been supplied yet. | |||||||