Qwen3

Qwen3-14B

Dense 14.8B-parameter Qwen3 checkpoint with thinking and non-thinking modes. The pinned configuration uses BF16 and a 40,960-token allocation; the publisher distinguishes native context from optional YaRN extension. Official vLLM recipe metadata lists a 34 GB BF16 minimum but reports Intel Xeon 6 CPU validation, so GPU fit, context memory and aggregate output capacity remain unverified.

Profile checked · Official model page ↗

Find rentals for this model →
Parameters
14.8B
Context
40960
BF16 weights + cacheOverhead additional
Estimate unavailable

Workload memory estimate

Parallel active users means one simultaneous request per user. Include input and expected output in the token allowance; idle users do not count. This sizes memory, with response speed still unverified.

Pinned checkpoint and cache configuration have not been supplied.

This is a single-GPU model-and-cache scenario. Distributed model sizing is not evaluated here. Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.

Find rentals for this model →

Supported deployment evidence

The finder evaluates exact, freshly documented rental configurations. Original BF16 is the default; quantized alternatives require opt-in. A profile describes sourced compatibility and memory allocation, with response speed and live stock assessed separately.

Sizing evidence is missing for deployment-aware search. An empty result does not establish that this model cannot run.

Hardware evidence

Benchmarked GPU profiles

Benchmark GPU / countArchitectureVRAM per GPUMatching on-demand instance
No benchmarked GPU profiles are available yet.

Measured performance

Inference benchmarks

Model GPU setup Engine Quant. Workload Reported tok/s Evidence Max-load capacity scenarioUSD / million output tokens
No benchmark records have been supplied yet.