Qwen3.5

Qwen3.5-9B

Publisher-labelled 9B language model with a vision encoder. The pinned post-trained checkpoint combines 24 Gated DeltaNet layers with eight full-attention layers. Text-only vLLM serving can omit the vision encoder. The model thinks by default. Full-context and concurrency memory need a separate hybrid-cache sizing method; no minimum GPU-memory claim or reachable throughput is asserted here.

Profile checked · Official model page ↗

Find rentals for this model →
Parameters
9B
Context
262144
BF16 weights + cacheOverhead additional
Estimate unavailable

Workload memory estimate

Parallel active users means one simultaneous request per user. Include input and expected output in the token allowance; idle users do not count. This sizes memory, with response speed still unverified.

This checkpoint's cache architecture needs a separate sizing method.

This is a single-GPU model-and-cache scenario. Distributed model sizing is not evaluated here. Runtime and allocator overhead are additional. This subtotal does not guarantee GPU compatibility. Input and output both contribute to cached tokens. No offloading, shared prefixes or cache quantization is assumed.

Find rentals for this model →

Supported deployment evidence

The finder evaluates exact, freshly documented rental configurations. Original BF16 is the default; quantized alternatives require opt-in. A profile describes sourced compatibility and memory allocation, with response speed and live stock assessed separately.

Sizing evidence is missing for deployment-aware search. An empty result does not establish that this model cannot run.

Hardware evidence

Benchmarked GPU profiles

Benchmark GPU / countArchitectureVRAM per GPUMatching on-demand instance
No benchmarked GPU profiles are available yet.

Measured performance

Inference benchmarks

Model GPU setup Engine Quant. Workload Reported tok/s Evidence Max-load capacity scenarioUSD / million output tokens
No benchmark records have been supplied yet.