OpenAI gpt-oss

gpt-oss-20b

Open-weight reasoning MoE with approximately 21B total and 3.6B active parameters. The original release uses native MXFP4 expert weights and the Harmony format. OpenAI describes operation within 16 GB for the released quantized model; serving context and concurrency require separate capacity evidence.

Profile checked · Official model page ↗

Find rentals for this model →
Parameters
21B
Context
131072
BF16 weights + cacheOverhead additional
Estimate unavailable

Workload memory estimate

Parallel active users means one simultaneous request per user. Include input and expected output in the token allowance; idle users do not count. This sizes memory, with response speed still unverified.

Pinned checkpoint and cache configuration have not been supplied.

This is a single-GPU model-and-cache scenario. Distributed model sizing is not evaluated here. Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.

Find rentals for this model →

Supported deployment evidence

The finder evaluates exact, freshly documented rental configurations. Original BF16 is the default; quantized alternatives require opt-in. A profile describes sourced compatibility and memory allocation, with response speed and live stock assessed separately.

Sizing evidence is missing for deployment-aware search. An empty result does not establish that this model cannot run.

Hardware evidence

Benchmarked GPU profiles

Benchmark GPU / countArchitectureVRAM per GPUMatching on-demand instance
No benchmarked GPU profiles are available yet.

Measured performance

Inference benchmarks

Model GPU setup Engine Quant. Workload Reported tok/s Evidence Max-load capacity scenarioUSD / million output tokens
No benchmark records have been supplied yet.