Qwen3
Qwen3-30B-A3B
Qwen3 checkpoint with sourced BF16 weights and full-attention cache configuration. MoE sizing includes all resident experts, not just active parameters.
Profile checked · Official model page ↗
- Parameters
- 30.5B
- Context
- 40960
- Min. VRAM
- —
Workload memory estimate
- Checkpoint weights
- 56.87 GiB
- KV cache
- 0.75 GiB
- Weights + cache
- 57.62 GiB
BF16 weights / BF16 cache · 8192 cached tokens per sequence · 1 resident sequence.
Pinned checkpoint ↗ · ad44e777bcd18fa416d9da3bd8f70d33ebb85d39 · huggingface.co: config.json ↗ · huggingface.co: model.safetensors.index.json ↗
Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility or pooled multi-GPU memory. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.
Hardware evidence
Benchmarked GPU profiles
| Benchmark GPU | Architecture | VRAM | Rental observed from |
|---|---|---|---|
| No benchmarked GPU profiles are available yet. | |||
Measured performance
Inference benchmarks
| Model | GPU setup | Engine | Quant. | Workload | Reported tok/s | Evidence | Max-load capacity scenarioUSD / million output tokens |
|---|---|---|---|---|---|---|---|
| No benchmark records have been supplied yet. | |||||||