OpenAI gpt-oss
gpt-oss-20b
Open-weight reasoning MoE with approximately 21B total and 3.6B active parameters. The original release uses native MXFP4 expert weights and the Harmony format. OpenAI describes operation within 16 GB for the released quantized model; serving context and concurrency require separate capacity evidence.
Profile checked · Official model page ↗
Find rentals for this model →- Parameters
- 21B
- Context
- 131072
- BF16 weights + cacheOverhead additional
- Estimate unavailable
Workload memory estimate
Parallel active users means one simultaneous request per user. Include input and expected output in the token allowance; idle users do not count. This sizes memory, with response speed still unverified.
Pinned checkpoint and cache configuration have not been supplied.
This is a single-GPU model-and-cache scenario. Distributed model sizing is not evaluated here. Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.
Find rentals for this model →Supported deployment evidence
The finder evaluates exact, freshly documented rental configurations. Original BF16 is the default; quantized alternatives require opt-in. A profile describes sourced compatibility and memory allocation, with response speed and live stock assessed separately.
Sizing evidence is missing for deployment-aware search. An empty result does not establish that this model cannot run.
Hardware evidence
Benchmarked GPU profiles
| Benchmark GPU / count | Architecture | VRAM per GPU | Matching on-demand instance |
|---|---|---|---|
| No benchmarked GPU profiles are available yet. | |||
Measured performance
Inference benchmarks
| Model | GPU setup | Engine | Quant. | Workload | Reported tok/s | Evidence | Max-load capacity scenarioUSD / million output tokens |
|---|---|---|---|---|---|---|---|
| No benchmark records have been supplied yet. | |||||||