For Qwen3-32B, the modeled API bill reaches the allocation budget around 53.72 million requests per 30 days, at 4,000 input and 500 output tokens per request. That is 26.86 billion output tokens, plus their associated input. But this crossover needs 10,362.32 average output tokens/second, above the published 3,462.00 output tokens/second for the matched allocation. That reference cannot establish a saving at the crossover. Investigate a better measured deployment before moving this traffic from the API.
Qwen3-32B can spend generated tokens on a thinking section before the visible answer. Start with the input and output tokens in your own API usage records, then compare the live bill with a complete allocation: NVIDIA's measured vLLM service used four two-GPU workers together. A separate two-GPU TensorRT-LLM recipe has no measured aggregate output rate in these sources. The worked crossover is a spending signal; verify capacity, latency and answer quality before changing deployments.
Count the thinking work, not just the displayed answer
Qwen3-32B's model card distinguishes thinking and non-thinking modes. Its template can include a thinking section before the final answer, and the non-thinking setting changes that generation path. If an application displays only the final answer, its visible length is a poor proxy for generated output. Segment your API usage records by mode and task, recording the provider-reported billable input and output tokens rather than guessing a billing rule for internal reasoning. The published serving example fixes each request at 4,000 input and 500 output tokens; a workload with long reasoning traces has a different shape and needs its own throughput test. Feed representative usage ratios into the live comparison, then check that the selected mode still meets your application's quality requirements.
Evidence: model source; model source; benchmark source.
The original memory budget and the FP8 service answer different questions
The pinned original checkpoint contains 61.02 GiB of BF16 tensor weights. Its 64 layers, eight KV heads and 128-element head dimension imply 256 KiB of BF16 key-value cache per resident token. Four sequences with 4,096 resident tokens each add 4.00 GiB of modeled cache, making 65.02 GiB before framework buffers. Four sequences at the catalog's 40,960-token allocation add 40.00 GiB of cache, taking the subtotal to 101.02 GiB. Qwen's card distinguishes its native context from optional YaRN extension, and explains why the default file allocates additional output room. Do not transfer those BF16 totals to NVIDIA's quantized service: its walkthrough uses an FP8 checkpoint and FP8 cache, while its measured example caps serving length at 6,000 tokens. Test memory and latency at the context lengths your own traffic actually uses.
Evidence: model source; architecture source; architecture source; architecture source; benchmark source.
The measured reference consumes eight H200 GPUs
NVIDIA's validation table reports 3,462 total output tokens per second for Qwen3-32B-FP8 on an eight-H200-SXM walkthrough. Four aggregated vLLM tensor-parallel-two workers sit behind one frontend; they are parts of the eight-GPU allocation, so multiplying the reported system result by four would count the same workers again. The AIPerf example holds prompts at 4,000 tokens, responses at 500, total concurrency at 56, and turns prefix caching off; the guide reports average time to first token, not a tail-latency guarantee. Its run date, measured checkpoint revision and exact measurement build are undisclosed. Reproduce that full topology if using this result as a capacity reference. NVIDIA also documents a separate two-GPU TensorRT-LLM recipe at concurrency four, but supplies no measured aggregate output rate for it here; benchmark that configuration separately before comparing its rent with the API.
Evidence: architecture source; benchmark source; serving source.
The deployment test that could change the decision
For each thinking mode, replay representative arrival bursts as well as steady traffic against the exact candidate checkpoint. Record the aggregate output rate, request-level latency distribution, cache use, failure rate and answer quality at your actual input and output lengths. The source's average latency and fixed-length validation cannot establish those properties for your users. The live equation below compares a full eight-GPU allocation against the current exact API endpoint; it also checks whether the required calendar-average output rate exceeds the published reference. If it does, first find a measured configuration that can carry the traffic, or a different priced allocation with its own matching benchmark. A lower modeled bill alone is not a deployment result.
Evidence: model source; benchmark source; serving source.
The current worked traffic calculation
The comparison uses DeepInfra's exact endpoint Qwen/Qwen3-32B (FP8), at 0.08 USD/million input tokens and 0.28 USD/million output tokens. The rental candidate is Crusoe Cloud, h200-141gb-sxm-ib.8x, eu-iceland1-a: 8 × H200 SXM 141 GB, at a whole-instance 34.320000 USD/hour. This is the least expensive qualifying comparison in our catalog, not a whole-market claim. Confirm current stock and final terms.
Period budget = 34.320000 USD/hour × 720 hours × 1 replica(s) + 0.00 USD
= 24,710.40 USD
API USD/million output, including input = 0.28 + (4000 / 500) × 0.08
= 0.92000000
Output-token crossover = 24,710.40 / 0.92000000 × 1,000,000
≈ 26,859,130,435 output tokens
Requests = output tokens / 500 ≈ 53,718,261| Requests per period | API bill, USD | Allocation budget, USD | Bill and throughput evidence |
|---|---|---|---|
| 100,000 | 46.00 | 24,710.40 | API lower bill |
| 1,000,000 | 460.00 | 24,710.40 | API lower bill |
| 10,000,000 | 4,600.00 | 24,710.40 | API lower bill |
| 42,974,608 | 19,768.32 | 24,710.40 | API lower bill |
| 64,461,913 | 29,652.48 | 24,710.40 | Allocation bill lower; above reference |
The crossover needs 10,362.32 output tokens/second averaged over the calendar period. The published FP8 reference is 3,462.00 output tokens/second per 8-GPU allocation, on 8 H200 SXM GPUs (walkthrough system h200_sxm), four TP2 workers behind one frontend, using Qwen/Qwen3-32B-FP8 and vLLM with NVIDIA Dynamo. The required rate exceeds the aggregate reference; evidence of a reachable saving is missing. This is not measured rental throughput, a latency guarantee, a peak ceiling, or proof that the API serves identical weights.
If the service sustained the published output rate continuously for 30 days at this request shape, it would produce 8.97 billion output tokens and the modeled API bill would be 8,255.62 USD, against the 24,710.40 USD allocation budget. This extrapolates a finite published test; it is not measured month-long capacity or an upper throughput bound.
API price checked 2026-09-30T06:16:25.505382+00:00; rental quote checked 2026-09-30T07:53:09.362914+00:00; published benchmark. Calculation generated 2026-09-30T19:31:14.978636+00:00. Listed compute excludes additional storage, traffic and operations unless included in the assumptions.
What changing request length would require
Only the 4,000-input / 500-output-token shape has a reviewed aggregate output reference here. A different input/output ratio needs a matching throughput measurement before this deployment comparison can claim capacity or savings for it.
Include the service you actually need
| Scenario | Period budget, USD | Break-even requests, millions |
|---|---|---|
| Current assumptions | 24,710.40 | 53.72 |
| Illustrative 1,000 USD operating budget | 25,710.40 | 55.89 |
| Twice the allocated replicas | 49,420.80 | 107.44 |
The operating allowance is illustrative. Doubling replicas is a cost scenario, not a tested availability design. Above the spending crossover the allocation bill can be lower on paper, but this published reference is below the needed average output rate. Measure a configuration that meets your traffic and latency requirements before claiming a saving.
Self-hosting terms and evidence
Qwen's original checkpoint and the FP8 checkpoint's own license each state Apache License 2.0. Its grant covers use and modification, including operating a self-hosted service. Redistribution requires retaining the license and applicable notices; modifications distributed as files need a prominent change notice. Review both exact files and any separate software dependencies for the intended deployment.
Publisher terms. Analysis reviewed 2026-09-30T13:51:18.491038+00:00. Prices refresh independently; an unchanged source check does not become a new editorial review. This analysis uses published reference measurements. Noach Ark did not rent this allocation or run inference or quality tests for this article.