For gpt-oss-120b, an API bill around 1,231.20 USD per 30 days can fund the modeled candidate compute budget. At 4,096 billed input and 1,024 billed output tokens per request, that is about 3.8 million requests, or 3.9 billion output tokens plus their input. Near this spending line, investigate the candidate below. Moving traffic requires measured capacity, acceptable task quality and a complete operating budget.

For a gpt-oss-120b API user, the decision is when a reasoning workload becomes large and steady enough to justify investigating an H100 deployment. The live calculation below turns API spending into a traffic target for one H100 SXM 80 GB. The distinctive catch is memory: OpenAI describes a single-80-GB release, while the captured vLLM guide warns that a default one-H100 configuration can run out of memory. Use the spending line to decide when to investigate; use your request lengths, bursts and latency requirements to decide whether the candidate can serve you.

Capacity unverified

This article provides a financial investigation screen. It has no qualified aggregate output-throughput measurement for the candidate. The traffic figure does not establish a reachable saving, supported concurrency, response time or equivalence to the API.

117B stored parameters still matter when only 5.1B are active

OpenAI describes gpt-oss-120b as 117 billion total parameters with 5.1 billion active per token. Its configuration routes each token through four of 128 local experts. The smaller active count describes sparse computation; it does not turn the stored model into a 5.1B checkpoint.

The released expert weights use MXFP4. The configuration excludes attention, routers, embeddings and the output head from that quantization, so treating every parameter as four-bit storage is also a poor sizing shortcut. The publisher associates this released model with a single 80 GB GPU. That is why the candidate here is an H100-class device rather than the small-GPU experiment used for gpt-oss-20b.

Keep the comparison precise: the self-hosted candidate is the pinned native MXFP4 release, while the exact API endpoint is recorded as bfloat16. The price comparison does not prove matching weights, numerical behavior or task quality. Replay the same successful tasks before moving traffic.

Evidence: model source; architecture source.

The one-H100 memory claim comes with a configuration test

The captured official vLLM guide specifically warns that gpt-oss-120b on an H100 at tensor-parallel size one can run out of memory with default GPU-memory utilization and batched-token settings. It gives higher GPU-memory utilization and a lower batched-token limit as a configuration to try. This is documented guidance, not a tested deployment from Noach Ark or a guarantee for your runtime version.

Before a throughput test, record the runtime version, image, maximum request length and batching limits, then measure peak memory with your longest requests. The checkpoint has 36 layers, split evenly between full and sliding-window attention, and a 131,072-token model limit. That limit does not establish how many long conversations one H100 can hold.

If the candidate requires more GPUs to meet your memory or latency target, rerun the budget for the actual allocation. A one-device spending line cannot justify a larger deployment. The vLLM document is a changing branch; this article reviews the captured snapshot, whose date and hash are retained with the evidence.

Evidence: architecture source; serving source.

Use billed reasoning work to choose the traffic target

OpenAI positions this larger model for production reasoning use and exposes low, medium and high reasoning effort. DeepInfra documents reasoning tokens as output-billed tokens. Use the API usage totals, including reasoning, rather than the visible answer length; do not add a reasoning subtotal twice when completion usage already includes it.

The worked example below uses 4,096 billed input and 1,024 billed output tokens per request. It is an explicit investigation workload, not a prediction of what this model generates. Compare it with your usage logs and the longer-output sensitivity before interpreting the request count. Keep reasoning effort, tool turns, retries and task-success criteria fixed across the API and the candidate.

Also plot when those requests arrive. A calendar-average output target cannot describe a concentrated peak or the queue that builds behind long reasoning jobs. For this H100 test, replay both ordinary traffic and the busiest real interval; measure how long complete successful tasks take.

Evidence: model source; serving source; serving source.

A faster token counter is not a passed migration test

The vLLM guide distinguishes aggregate output-token throughput from combined input-plus-output throughput. Compare the required rate below with generated output over full request wall time. Its sample benchmark report contains placeholder values; it supplies no qualified capacity number for the pinned one-H100 candidate in this article.

The guide also recommends a streaming interval of 20 tokens for throughput-oriented configurations. A change to streaming cadence can affect what a user sees, so check time to first token and end-to-end latency alongside aggregate throughput. Record the actual concurrency and request lengths instead of carrying a different hardware or workload result into this decision.

OpenAI requires Harmony formatting. Include reasoning/final-answer separation and tool-call continuation in the same replay. Investigate near the live spending line, then move traffic only if memory, sustained output, peak latency and task success pass together with your complete operating budget. The missing measurement is the candidate service, not another multiplication of its active parameter count.

Evidence: model source; serving source.

The candidate to test

Test 1 × H100 SXM 80 GB (80 GB per GPU) with the pinned MXFP4 checkpoint. Its 80 GB publisher memory claim supports investigating this single-device candidate. It does not size your context, simultaneous requests or runtime buffers. The API is separately tracked as bfloat16. Review checkpoint/precision equivalence and task quality; this calculation assumes the same billed token workload, without proving that the candidate generates the same tokens.

Spending investigation threshold

The current API observation is DeepInfra's openai/gpt-oss-120b, bfloat16, at 0.037 USD/million input and 0.17 USD/million output tokens. The lowest qualifying quote in our catalog for this exact candidate is Vast.ai, Marketplace offer #50189137, Australia, AU, 1.71 USD/hour for the whole instance. It is listed on-demand hourly pricing with no minimum commitment; confirm stock and final charges. Stock status: listed. This is catalog coverage, not a whole-market comparison.

# Display values are rounded; totals use full stored rates.
Period budget ≈ 1.71 USD/hour × 720 hours × 1 replica(s) + 0.00 USD
              ≈ 1,231.20 USD
API USD/million output, including input = 0.17 + (4096 / 1024) × 0.037
                                       = 0.318
Spending investigation threshold = budget / API USD per million output × 1,000,000
                                ≈ 3.9 billion billed output tokens
Requests ≈ 3.8 million per 30 days
Model-specific traffic and cost calculation
Requests per period API bill, USD Candidate budget, USD Financial screen
100 thousand 32.56 1,231.20 Below candidate budget
1 million 325.63 1,231.20 Below candidate budget
3 million 976.90 1,231.20 Below candidate budget
4.5 million 1,465.34 1,231.20 Spend can fund a capacity test

At the spending threshold the workload demands 1,494 billed output tokens/second averaged over the calendar period, or 1,494 per replica if traffic is evenly shared. These are measurement targets, not measured GPU speeds. Replay the actual request lengths, bursts, reasoning settings and tool turns, then check aggregate output throughput, tail latency, failures and task success.

With the same 4,096 input tokens but 8,192 billed output tokens per request, the same budget corresponds to 800 thousand requests and 2,520 average output tokens/second. This longer-output sensitivity illustrates the effect of billed generation; it does not predict a reasoning setting's output length or the GPU's performance.

The period assumes 720 always-on hours and 1 replica(s). The explicit additional operating allowance is 0.00 USD. Include storage, network, taxes, monitoring, operational work and redundancy before comparing total service cost. When this allowance is zero the screen covers compute only. Extra replicas increase the bill; they are not a tested availability design.

API checked October 6, 2026 at 9:30 am EDT, collector source; rental checked October 6, 2026 at 11:00 am EDT, quote source. Recorded terms: On-demand compute billed for actual active runtime; no monthly commitment. Storage persists and remains billable until deleted; transfer and other charges are separate. Quote notes: API discovery sample, capped at 12/GPU model. Rentable at check; capacity can change. Compute only; storage and traffic extra. Prices and calculations refresh automatically. Figures are rounded for readability; calculations use the full stored precision. Stale, retired or changed-identity observations suppress this calculation.

Self-hosting terms and evidence

The pinned publisher checkpoint is distributed under Apache 2.0. Review and follow its conditions for use, modification and redistribution, including required notices and identification of distributed changes. Its accompanying USAGE_POLICY requires compliance with applicable law. These captured publisher terms have been reviewed for this self-hosting comparison; the model license does not establish compliance of a particular application.

Publisher terms. Analysis reviewed October 6, 2026. Noach Ark did not rent this allocation or run inference, capacity or quality tests for this article. Prices and financial calculations refresh independently of this editorial review.