For Qwen3.5-9B, an API bill around 316.80 USD per 30 days can fund the modeled candidate compute budget. At 2,048 billed input and 128 billed output tokens per request, that is about 1.4 million requests, or 180 million output tokens plus their input. Near this spending line, investigate the candidate below. Moving traffic requires measured capacity, acceptable task quality and a complete operating budget.

Qwen3.5-9B combines a vision encoder, hybrid attention and default thinking. For a text-only application, use the live spending line to decide when to investigate the original BF16 checkpoint on one L4. Define the text-only boundary, thinking mode and longest real prompts before testing; each can change what your deployment needs to serve.

Capacity unverified

This article provides a financial investigation screen. It has no qualified aggregate output-throughput measurement for the candidate. The traffic figure does not establish a reachable saving, supported concurrency, response time or equivalence to the API.

Text-only serving changes what you load

Qwen provides a text-only vLLM command using --language-model-only to omit the vision encoder and multimodal profiling. For an application that sends only text, test that configuration first. An application that reads images or video needs a separate deployment comparison. Keep the workload boundary explicit when measuring memory and throughput.

Evidence: model source.

The hybrid cache needs its own memory method

The pinned configuration has 32 layers: 24 linear-attention layers and eight full-attention layers, with four KV heads of dimension 256 in the latter. It also declares a native 262,144-token context. Multiplying a conventional KV-cache formula by all 32 layers would misdescribe this checkpoint; accounting only for the eight full-attention layers would omit the linear-attention state and other allocations. Our catalog records the hybrid sizing method as unsupported. Measure the memory of the actual runtime at your longest prompts and concurrency before selecting a device.

Evidence: architecture source.

Match the thinking mode before counting requests

Qwen3.5 thinks by default. Its local API example disables thinking through chat_template_kwargs with enable_thinking set to false; the publisher warns that /think and /nothink switches are unsupported. Fix this setting across your hosted and local comparison, then measure task success, generated output and response time. The publisher supplies separate reasoning and tool parsers; verify both for an application using tools.

Evidence: model source.

The candidate to test

Test 1 × L4 24 GB (24 GB per GPU) with the pinned BF16 checkpoint. Serving-engine guidance: the captured official vLLM recipe lists a 22 GB minimum for this model and precision, with vLLM 0.17.0 or newer. Serving recipe, captured October 5, 2026 at 9:43 am EDT; the reviewed snapshot is fixed by its capture hash, even if the linked document changes. Reported validation scope: The recipe records single-device Intel Arc Pro B60/B70 validation with the vLLM XPU image, an 8,192-token maximum model length and eager mode. It lists the L4 as candidate hardware but supplies no L4 measurement or tested weight revision. Its separate 262,144-token launch example is not a validated L4 capacity result. GPU fit unverified: this guidance identifies a candidate to investigate; it does not verify the chosen GPU, checkpoint revision, context length, simultaneous requests or runtime memory. The checkpoint is pinned for your test; the engine recipe names the model without identifying its tested checkpoint revision. The API is separately tracked as bfloat16. Verify fit, formatting, task quality and checkpoint/precision equivalence before moving traffic. The calculation assumes the same billed token workload, without proving that the candidate generates the same tokens.

Spending investigation threshold

The current API observation is DeepInfra's Qwen/Qwen3.5-9B, bfloat16, at 0.1 USD/million input and 0.15 USD/million output tokens. The lowest qualifying quote in our catalog for this exact candidate is Jarvislabs, 1x L4 24 GB, region not specified, 0.44 USD/hour for the whole instance. It is listed on-demand hourly pricing with no minimum commitment; confirm stock and final charges. Stock status: unverified. This is catalog coverage, not a whole-market comparison.

# Display values are rounded; totals use full stored rates.
Period budget ≈ 0.44 USD/hour × 720 hours × 1 replica(s) + 0.00 USD
              ≈ 316.80 USD
API USD/million output, including input = 0.15 + (2048 / 128) × 0.1
                                       = 1.75
Spending investigation threshold = budget / API USD per million output × 1,000,000
                                ≈ 180 million billed output tokens
Requests ≈ 1.4 million per 30 days
Model-specific traffic and cost calculation
Requests per period API bill, USD Candidate budget, USD Financial screen
100 thousand 22.40 316.80 Below candidate budget
1 million 224.00 316.80 Below candidate budget
1.1 million 246.40 316.80 Below candidate budget
1.7 million 380.80 316.80 Spend can fund a capacity test

At the spending threshold the workload demands 70 billed output tokens/second averaged over the calendar period, or 70 per replica if traffic is evenly shared. These are measurement targets, not measured GPU speeds. Replay the actual request lengths, bursts, reasoning settings and tool turns, then check aggregate output throughput, tail latency, failures and task success.

With the same 2,048 input tokens but 1,024 billed output tokens per request, the same budget corresponds to 880 thousand requests and 349 average output tokens/second. This longer-output sensitivity illustrates the effect of billed generation; it does not predict a reasoning setting's output length or the GPU's performance.

The period assumes 720 always-on hours and 1 replica(s). The explicit additional operating allowance is 0.00 USD. Include storage, network, taxes, monitoring, operational work and redundancy before comparing total service cost. When this allowance is zero the screen covers compute only. Extra replicas increase the bill; they are not a tested availability design.

API checked October 5, 2026 at 9:47 am EDT, collector source; rental checked October 5, 2026 at 9:09 am EDT, quote source. Recorded terms: On-demand; no monthly commitment. Quote notes: USD outside India; storage and tax additional. Prices and calculations refresh automatically. Figures are rounded for readability; calculations use the full stored precision. Stale, retired or changed-identity observations suppress this calculation.

Self-hosting terms and evidence

The pinned model is distributed under Apache 2.0, which permits use, reproduction and distribution subject to its conditions. Preserve required license and attribution notices, mark distributed modifications and review patent and trademark conditions for the actual deployment.

Publisher terms. Analysis reviewed October 5, 2026. Noach Ark did not rent this allocation or run inference, capacity or quality tests for this article. Prices and financial calculations refresh independently of this editorial review.