For gpt-oss-20b, an API bill around 316.80 USD per 30 days can fund the modeled candidate compute budget. At 2,048 billed input and 128 billed output tokens per request, that is about 3.99 million requests, or 510.97 million output tokens plus their input. Near this spending line, investigate the candidate below. Moving traffic requires measured capacity, acceptable task quality and a complete operating budget.

For gpt-oss-20b, the useful question is whether your complete reasoning-task bill can fund a small-GPU experiment. Its native MXFP4 checkpoint makes that experiment plausible, but the active parameter count does not size the service. Count all billed generation and tool turns, keep reasoning effort fixed, and test the actual checkpoint before treating lower compute spending as a saving.

Capacity unverified

This article provides a financial investigation screen. It has no qualified aggregate output-throughput measurement for the candidate. The traffic figure does not establish a reachable saving, supported concurrency, response time or equivalence to the API.

A short answer can still be an expensive reasoning task

gpt-oss-20b supports low, medium and high reasoning effort. DeepInfra documents reasoning tokens as output-billed tokens. Count the completion usage reported by the API, including reasoning, instead of measuring only the final answer. Do not add a reasoning subtotal twice when it is already included in completion usage.

For each completed task, record billed input and output, reasoning setting, model-call count, retries and task success. An agent can make several calls and send tool results back as fresh input. Compare the complete successful task on the API and the candidate. Lower cost per token has little value if more retries or weaker answers erase it.

The calculation uses an illustrative request shape, not a prediction of this model's generation length. The longer-output sensitivity below shows how billed generation moves the request-count spending line. Both token volume and the capacity measurement target change; the reasoning setting alone does not supply either number.

Evidence: model source; serving source.

Native MXFP4 makes a small-GPU experiment plausible

OpenAI describes gpt-oss-20b as about 21 billion total parameters with 3.6 billion active per token. The checkpoint selects four of its 32 experts per token. That sparsity reduces the work performed for a token; the server still needs the stored checkpoint.

The released expert weights were post-trained in MXFP4, and the publisher reports running the model within 16 GB. This supports investigating the original checkpoint on a device with more memory, rather than estimating its footprint as a dense BF16 model. It does not establish context capacity, simultaneous requests or runtime headroom.

The tracked API observation is labelled bfloat16; the candidate is native MXFP4. The spending calculation can compare those bills, but cannot establish weight equivalence or equal output quality. Keep the precision labels visible and replay the same tasks before considering a switch.

Evidence: model source; architecture source.

Mixed attention makes the context limit a poor capacity shortcut

The checkpoint has 24 attention layers combining sliding-window and full attention. Its configuration sets a 128-token sliding window and a 131,072-token model limit. A sliding window does not apply to every layer, so the complete cache does not become a fixed 128-token allocation.

The model limit is a capability of the checkpoint, not the number of long conversations that one L4 can serve. Test the longest prompts and simultaneous requests you actually need, alongside short requests. A deployment that fits the model at a small context may fail your production memory or response-time requirements. This article deliberately provides no BF16 full-attention cache estimate for the native MXFP4, mixed-attention checkpoint.

Evidence: architecture source.

Harmony handling belongs in the migration test

OpenAI says gpt-oss must use its Harmony response format to work correctly. A server accepting an OpenAI-compatible HTTP request does not by itself establish correct formatting. Verify the chat template, separation of reasoning from the final answer, tool-call parsing and continuation after tool results.

Start with the pinned original checkpoint below and a serving runtime that supports its MXFP4 weights on the chosen GPU. Keep reasoning effort and tool behavior fixed across your API and candidate runs. Count all generated tokens, measure aggregate output throughput over full request wall time, and check tail latency and task success. No qualified throughput result is attached to this candidate, so the required average rate below is a test target. Investigate near the spending line; move traffic only after these measurements and your complete operating budget support the decision.

Evidence: model source.

The candidate to test

Test 1 × L4 24 GB (24 GB per GPU) with the pinned MXFP4 checkpoint. Its 16 GB publisher memory claim supports investigating this single-device candidate. It does not size your context, simultaneous requests or runtime buffers. The API is separately tracked as bfloat16. Review checkpoint/precision equivalence and task quality; this calculation assumes the same billed token workload, without proving that the candidate generates the same tokens.

Spending investigation threshold

The current API observation is DeepInfra's openai/gpt-oss-20b, bfloat16, at 0.03000000 USD/million input and 0.14000000 USD/million output tokens. The lowest qualifying quote in our catalog for this exact candidate is Jarvislabs, 1x L4 24 GB, region not specified, 0.440000000000000000 USD/hour for the whole instance. It is listed on-demand hourly pricing with no minimum commitment; confirm stock and final charges. Stock status: unverified. This is catalog coverage, not a whole-market comparison.

Period budget = 0.440000000000000000 USD/hour × 720 hours × 1 replica(s) + 0.00 USD
              = 316.800000000000000000 USD
API USD/million output, including input = 0.14000000 + (2048 / 128) × 0.03000000
                                       = 0.62000000
Spending investigation threshold = budget / API USD per million output × 1,000,000
                                ≈ 510,967,742 billed output tokens
Requests ≈ 3,991,935 per 30 days
Model-specific traffic and cost calculation
Requests per period API bill, USD Candidate budget, USD Financial screen
100,000 7.94 316.80 Below candidate budget
1,000,000 79.36 316.80 Below candidate budget
3,193,548 253.44 316.80 Below candidate budget
4,790,322 380.16 316.80 Spend can fund a capacity test

At the spending threshold the workload demands 197.13 billed output tokens/second averaged over the calendar period, or 197.13 per replica if traffic is evenly shared. These are measurement targets, not measured GPU speeds. Replay the actual request lengths, bursts, reasoning settings and tool turns, then check aggregate output throughput, tail latency, failures and task success.

With the same 2,048 input tokens but 1,024 billed output tokens per request, the same budget corresponds to 1.55 million requests and 611.11 average output tokens/second. This longer-output sensitivity illustrates the effect of billed generation; it does not predict a reasoning setting's output length or the GPU's performance.

The period assumes 720 always-on hours and 1 replica(s). The explicit additional operating allowance is 0.00 USD. Include storage, network, taxes, monitoring, operational work and redundancy before comparing total service cost. When this allowance is zero the screen covers compute only. Extra replicas increase the bill; they are not a tested availability design.

API checked 2026-10-04T06:15:17.521765+00:00, collector source; rental checked 2026-10-02T13:05:26.985627+00:00, quote source. Recorded terms: On-demand; no monthly commitment. Quote notes: USD outside India; storage and tax additional. Calculated 2026-10-04T07:57:35.510988+00:00. Stale, retired or changed-identity observations suppress this calculation.

Self-hosting terms and evidence

The pinned checkpoint is distributed under Apache 2.0. Follow its conditions when using, modifying or redistributing the model, including preservation of required notices and identification of changes when distributing modified material. The accompanying USAGE_POLICY requires compliance with applicable law. Review these publisher files for your actual deployment and distribution; the model license does not establish application compliance.

Publisher terms. Analysis reviewed 2026-10-04T06:30:13.260285+00:00. Noach Ark did not rent this allocation or run inference, capacity or quality tests for this article. Prices and financial calculations refresh independently of this editorial review.