Qwen3

Qwen3-30B-A3B

Qwen3 checkpoint with sourced BF16 weights and full-attention cache configuration. MoE sizing includes all resident experts, not just active parameters.

Profile checked · Official model page ↗

Parameters
30.5B
Context
40960
Min. VRAM
—

Workload memory estimate

Checkpoint weights
56.87 GiB
KV cache
0.75 GiB
Weights + cache
57.62 GiB

BF16 weights / BF16 cache · 8192 cached tokens per sequence · 1 resident sequence.

Pinned checkpoint ↗ · ad44e777bcd18fa416d9da3bd8f70d33ebb85d39 · huggingface.co: config.json ↗ · huggingface.co: model.safetensors.index.json ↗

Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility or pooled multi-GPU memory. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.

Hardware evidence

Benchmarked GPU profiles

Benchmark GPUArchitectureVRAMRental observed from
No benchmarked GPU profiles are available yet.

Measured performance

Inference benchmarks

Model GPU setup Engine Quant. Workload Reported tok/s Evidence Max-load capacity scenarioUSD / million output tokens
No benchmark records have been supplied yet.