For Qwen3.5-35B-A3B, an API bill around 2,880.00 USD per 30 days can fund the modeled candidate compute budget. At 4,096 billed input and 1,024 billed output tokens per request, that is about 1.8 million requests, or 1.8 billion output tokens plus their input. Near this spending line, investigate the candidate below. Moving traffic requires measured capacity, acceptable task quality and a complete operating budget.
Your Qwen3.5-35B-A3B traffic warrants investigating self-hosting when its billed usage approaches the live compute budget below. This comparison tests a specific option: the original BF16 checkpoint on one H200 NVL 141 GB, against the exact FP8 API endpoint. The model activates about 3B parameters per token, but that number does not size its stored experts or certify a cheap deployment. Use the spending line to decide when a test deserves attention. The test must establish GPU fit, successful-task quality and capacity under your arrival pattern before you move traffic.
Capacity unverified
This article provides a financial investigation screen. It has no qualified aggregate output-throughput measurement for the candidate. The traffic figure does not establish a reachable saving, supported concurrency, response time or equivalence to the API.
3B active still leaves 256 experts in your memory plan
Qwen labels this model 35B total and 3B active. Its pinned configuration selects eight of 256 routed experts per token and also includes a shared expert. Sparse computation therefore changes the work performed for a token, while the selected original checkpoint still contains the expert weights. An active-parameter count is not a device-memory requirement.
The captured official vLLM recipe lists an 84 GB minimum for the default BF16 variant. That guidance does not qualify a one-80-GB GPU comparison. It gives a reason to investigate the 141 GB H200 candidate used here. Its explicit GPU validation, however, is four Intel Arc Pro B60/B70 cards at an 8,192-token maximum model length, alongside Intel CPU support. The guide supplies no measured one-H200 result or tested publisher weight revision. GPU fit remains unverified.
The same recipe assigns 42 GB guidance to a separately named official FP8 checkpoint. That is an alternative experiment with different weights and precision; it does not reduce this BF16 candidate budget. The tracked API is FP8 too, so replay successful tasks and compare generation usage before treating these two services as equivalent.
Evidence: model source; architecture source; serving source.
Repeated prompts need a hybrid-state test
This checkpoint has 40 layers: 30 linear-attention layers and ten full-attention layers. Its configuration also sets the recurrent Mamba state to float32. A uniform BF16 KV-cache calculation would miss that mixture; a calculation using only the full-attention layers would omit the recurrent state and runtime allocations. No conventional full-attention memory estimate is assigned to this candidate.
That architecture matters when your API workload repeatedly sends a long system prompt or conversation history. The captured vLLM recipe calls Mamba prefix caching experimental in its align mode. Do not assume a repeated prefix has the same serving effect as in a conventional attention model. Replay requests with unique prefixes, repeated prefixes and multi-turn histories, recording the cache policy and memory actually allocated.
The guide separately discusses CUDA-graph/Mamba cache-size errors and suggests reducing the maximum CUDA-graph capture size. Treat this as a version-specific configuration to investigate, not a measured H200 solution. First establish that the longest real requests fit; then measure aggregate output and tail latency at the concurrency you need. The native 262,144-token limit does not establish the number of conversations one H200 can hold.
Evidence: architecture source; serving source.
MTP is a second test configuration, not a speed multiplier
The publisher includes multi-token prediction in this model. Its captured vLLM example uses qwen3_next_mtp with two speculative tokens, while the captured official engine recipe uses mtp with one. These are documented commands from different source snapshots, not a measured comparison. Pin a compatible runtime and its actual speculative settings instead of combining flags from the two examples.
Run the ordinary serving configuration first, then test MTP on the same request replay. Count final generated output over full request wall time, rather than internal draft proposals or a combined input-plus-output counter. Compare peak memory, time to first token, completion latency and successful tasks. No sourced multiplier is attached to the H200 candidate, and MTP cannot supply a capacity number by itself.
Keep the thinking setting fixed across those tests. Qwen3.5 thinks by default and does not officially support the older /think and /nothink prompt switches; its example disables thinking through enable_thinking=False in chat-template parameters. DeepInfra bills reasoning tokens as output. Use reported billed output, including reasoning without double-counting it, to choose your traffic shape. The live example uses 4,096 input and 1,024 output tokens; it is a measurement workload, not a predicted generation length.
Evidence: model source; serving source; serving source.
Check which Qwen service your bill actually represents
The publisher describes Qwen3.5-Flash as a corresponding hosted service with additional production features, including a default 1M context and built-in tools. This article instead pins DeepInfra's exact Qwen/Qwen3.5-35B-A3B endpoint. A Qwen3.5-Flash bill cannot be inserted into this endpoint's live token calculation without reviewing its own rates, billed usage and included features.
For a migration, list the tools and context behavior supplied by your current service. The downloadable checkpoint declares a native 262,144-token context and a separate extension configuration; it does not bundle the hosted service's tool execution or infrastructure. Include that work in the candidate operating plan.
The financial example here concerns billed text input and output. This is a vision-language checkpoint, but image/video preprocessing and media-heavy requests are not measured by that text replay. Test those separately if you use them. Removing the vision encoder changes the serving configuration and its memory behavior, so it also needs a separate fit and capacity test. The purpose of the H200 spending screen is to choose a concrete investigation, not to claim that every Qwen3.5 service or workload is interchangeable.
Evidence: model source; architecture source; serving source.
The candidate to test
Test 1 × H200 NVL (141 GB per GPU) with the pinned BF16 checkpoint. Serving-engine guidance: the captured official vLLM recipe lists a 84 GB minimum for this model and precision, with vLLM 0.17.0 or newer. Serving recipe, captured October 9, 2026 at 9:30 am EDT; the reviewed snapshot is fixed by its capture hash, even if the linked document changes. Reported validation scope: The captured recipe marks Intel Xeon 6 CPU and Arc Pro B60/B70 hardware verified. Its explicit GPU example uses four Intel Arc Pro B60/B70 cards with the vLLM XPU image, tensor parallel size four, eager mode and an 8,192-token maximum model length. BF16 prerequisites list one H200 without naming its variant, while the launch example uses two H200s at 262,144 tokens. Neither example provides a measured one-H200 NVL workload or tested publisher weight revision. GPU fit for this single-device candidate remains unverified. GPU fit unverified: this guidance identifies a candidate to investigate; it does not verify the chosen GPU, checkpoint revision, context length, simultaneous requests or runtime memory. The checkpoint is pinned for your test; the engine recipe names the model without identifying its tested checkpoint revision. The API is separately tracked as fp8. Verify fit, formatting, task quality and checkpoint/precision equivalence before moving traffic. The calculation assumes the same billed token workload, without proving that the candidate generates the same tokens.
Spending investigation threshold
The current API observation is DeepInfra's Qwen/Qwen3.5-35B-A3B, fp8, at 0.14 USD/million input and 1 USD/million output tokens. The lowest qualifying quote in our catalog for this exact candidate is Vast.ai, Marketplace offer #45172707, Japan, JP, 4 USD/hour for the whole instance. It is listed on-demand hourly pricing with no minimum commitment; confirm stock and final charges. Stock status: listed. This is catalog coverage, not a whole-market comparison.
# Display values are rounded; totals use full stored rates.
Period budget ≈ 4 USD/hour × 720 hours × 1 replica(s) + 0.00 USD
≈ 2,880.00 USD
API USD/million output, including input = 1 + (4096 / 1024) × 0.14
= 1.56
Spending investigation threshold = budget / API USD per million output × 1,000,000
≈ 1.8 billion billed output tokens
Requests ≈ 1.8 million per 30 days| Requests per period | API bill, USD | Candidate budget, USD | Financial screen |
|---|---|---|---|
| 100 thousand | 159.74 | 2,880.00 | Below candidate budget |
| 1 million | 1,597.44 | 2,880.00 | Below candidate budget |
| 1.4 million | 2,236.42 | 2,880.00 | Below candidate budget |
| 2.2 million | 3,514.37 | 2,880.00 | Spend can fund a capacity test |
At the spending threshold the workload demands 712 billed output tokens/second averaged over the calendar period, or 712 per replica if traffic is evenly shared. These are measurement targets, not measured GPU speeds. Replay the actual request lengths, bursts, reasoning settings and tool turns, then check aggregate output throughput, tail latency, failures and task success.
With the same 4,096 input tokens but 8,192 billed output tokens per request, the same budget corresponds to 330 thousand requests and 1,038 average output tokens/second. This longer-output sensitivity illustrates the effect of billed generation; it does not predict a reasoning setting's output length or the GPU's performance.
The period assumes 720 always-on hours and 1 replica(s). The explicit additional operating allowance is 0.00 USD. Include storage, network, taxes, monitoring, operational work and redundancy before comparing total service cost. When this allowance is zero the screen covers compute only. Extra replicas increase the bill; they are not a tested availability design.
API checked October 9, 2026 at 9:34 am EDT, collector source; rental checked October 9, 2026 at 3:01 pm EDT, quote source. Recorded terms: On-demand compute billed for actual active runtime; no monthly commitment. Storage persists and remains billable until deleted; transfer and other charges are separate. Quote notes: API discovery sample, capped at 12/GPU model. Rentable at check; capacity can change. Compute only; storage and traffic extra. Prices and calculations refresh automatically. Figures are rounded for readability; calculations use the full stored precision. Stale, retired or changed-identity observations suppress this calculation.
Self-hosting terms and evidence
The pinned Qwen publisher checkpoint is licensed under Apache 2.0. The captured terms have been reviewed for this self-hosting comparison. Follow the conditions for use, modification and redistribution, including giving recipients the license, preserving applicable notices and identifying modified files when distributing changes. The license does not establish the suitability or legal compliance of a particular application.
Publisher terms. Analysis reviewed October 9, 2026. Noach Ark did not rent this allocation or run inference, capacity or quality tests for this article. Prices and financial calculations refresh independently of this editorial review.