For Gemma 4 31B IT, an API bill around 2,235.84 USD per 30 days can fund the modeled candidate compute budget. At 4,096 billed input and 1,024 billed output tokens per request, that is about 1.8 million requests, or 1.9 billion output tokens plus their input. Near this spending line, investigate the candidate below. Moving traffic requires measured capacity, acceptable task quality and a complete operating budget.

Your Gemma 4 31B traffic warrants investigating self-hosting when its API spending approaches the live candidate budget below. The concrete experiment is the original BF16 checkpoint on one H100 SXM 80 GB, compared with DeepInfra's exact FP8 endpoint. For this model, image detail and conversation formatting belong in the test alongside token volume: a financial threshold cannot tell you whether a dense vision model will fit your requests or complete the same tasks. Use the spending line to decide when to investigate, then establish fit, quality and serving capacity before moving traffic.

Capacity unverified

This article provides a financial investigation screen. It has no qualified aggregate output-throughput measurement for the candidate. The traffic figure does not establish a reachable saving, supported concurrency, response time or equivalence to the API.

The 31B launch is not the guide's single-GPU quick start

Gemma 4 includes several different deployment sizes. This article pins the 31B dense instruction checkpoint, not the 26B sparse model or the E4B model. Its publisher configuration explicitly disables the MoE block; there is no smaller active-expert count to use as a shortcut for this checkpoint's memory needs.

The captured official vLLM YAML lists a 75 GB minimum for its default BF16 variant and a minimum engine version of 0.19.1. That admits an 80 GB H100 SXM as a financial candidate. Its actual 31B BF16 launch example uses two A100/H100 GPUs with tensor parallelism two and a 32,768-token maximum model length. Its full-featured 31B example also uses two GPUs, at 16,384 tokens. The guide's single-GPU quick start serves E4B, and its one-A100/H100 example serves the 26B sparse checkpoint. Neither validates this 31B candidate.

Treat 75 GB as captured serving guidance, not a measured allocation on one H100 SXM. The guide marks H100 hardware verified without supplying a single-SXM workload result or tested publisher weight revision. GPU fit remains unverified. The FP8 API is a separate precision: replay successful tasks before treating its output and the original BF16 model as interchangeable.

Evidence: model source; architecture source; serving source.

Image detail changes the experiment; audio belongs to another variant

The publisher gives Gemma 4 five visual token budgets: 70, 140, 280, 560 and 1,120. It recommends smaller budgets for classification, captioning or video understanding, and larger budgets for OCR, document parsing and small text. These are different workloads. A captioning replay cannot establish that the same allocation handles your detailed document requests at the required latency.

For a document workflow, replay the pages and crops that contain the smallest text you actually need to read. Record the processor settings, visual budget, input accounting, task success and peak memory. For video, preserve the real frame count and sampling policy; the publisher describes up to 60 seconds at one frame per second. Do not equate a frame-heavy request with the text-token example in the financial block. That example covers reported billed text input and output; it provides no measured image or video capacity.

There is also a model-specific trap in the serving guide: its family description and full-featured 31B launch mention audio. The pinned publisher configuration has audio_config set to null, and the model card assigns audio input to E2B, E4B and 12B. This 31B checkpoint has a vision encoder and supports text, images and video frames; it does not supply the audio encoder described for those other variants. An audio workflow needs separate model evidence and economics. Copying the family launch command would not establish that feature.

Evidence: model source; architecture source; serving source.

The long-context test has two different attention paths

The pinned text configuration has 60 layers: 50 sliding-attention layers with a 1,024-token window and ten full-attention layers. It specifies different local and global head dimensions, different key/value head counts, and shared keys and values for global attention. This mixture makes a uniform full-attention cache formula a poor basis for sizing the candidate. This spending screen assigns no estimated BF16 cache or promised conversation count.

The model declares a native 262,144-token context, while the captured 31B engine examples use much shorter maximum lengths. A model's context limit is not a tested memory budget for several simultaneous long conversations, particularly with a vision encoder and image inputs. The guide's 75 GB minimum does not settle the cache, processor, runtime-buffer or concurrency headroom.

Build the fit test from your actual combination of prompt length, image count and visual budget. Test one longest request first, then simultaneous requests and bursts, preserving the runtime settings. Measure memory allocated, time to first token, completion latency, failures and aggregate generated output over full request wall time. Count completed task output rather than a per-user decode rate or input-plus-output total. The live average-output figure is a target to measure against, not an H100 performance result.

Evidence: model source; architecture source; serving source.

A tool conversation needs more than an OpenAI-shaped request

The pinned publisher template requires tool_calls[].function.arguments to be a parsed JSON object. It raises an exception when given an argument string and explicitly instructs callers to deserialize it first. If your API client stores arguments as JSON strings, replaying that history directly through this template can fail before any capacity test begins. Normalize the history according to the actual template you serve, while preserving the tool names and arguments.

Thinking also has turn-specific rules. The publisher says to discard thoughts from ordinary previous assistant turns but preserve them for tool-call turns. Its template defaults enable_thinking to false and has a separate preserve_thinking gate. The card says the 31B model still emits an empty thought wrapper when thinking is disabled. Test the parser and the next turn, not just whether the first answer looks readable.

The official engine example uses the gemma4 reasoning and tool-call parsers together with an engine-provided tool chat template. That is a different template from the pinned publisher file; record which one your runtime actually uses rather than assuming the two histories are interchangeable. Replay complete tool loops, including a tool result and the next assistant turn. Track reported token usage, tool execution, retries and successful tasks alongside latency. A fast isolated answer that breaks the next tool turn does not qualify the migration.

Evidence: model source; architecture source; serving source.

The candidate to test

Test 1 × H100 SXM 80 GB (80 GB per GPU) with the pinned BF16 checkpoint. Serving-engine guidance: the captured official vLLM recipe lists a 75 GB minimum for this model and precision, with vLLM 0.19.1 or newer. Serving recipe, captured October 10, 2026 at 9:32 am EDT; the reviewed snapshot is fixed by its capture hash, even if the linked document changes. Reported validation scope: The captured YAML marks H100, MI300X, MI325X, MI355X, Trillium and Ironwood hardware verified. Its explicit 31B BF16 NVIDIA example uses two A100/H100 GPUs with tensor parallel size two and a 32,768-token maximum model length; the full-featured 31B example uses two GPUs at 16,384 tokens. The single-GPU quick start is E4B and the one-A100/H100 example is the different 26B sparse model. No measured one-H100 SXM workload or tested publisher weight revision is provided. GPU fit for this single-device candidate remains unverified. GPU fit unverified: this guidance identifies a candidate to investigate; it does not verify the chosen GPU, checkpoint revision, context length, simultaneous requests or runtime memory. The checkpoint is pinned for your test; the engine recipe names the model without identifying its tested checkpoint revision. The API is separately tracked as fp8. Verify fit, formatting, task quality and checkpoint/precision equivalence before moving traffic. The calculation assumes the same billed token workload, without proving that the candidate generates the same tokens.

Spending investigation threshold

The current API observation is DeepInfra's google/gemma-4-31B-it, fp8, at 0.2 USD/million input and 0.4 USD/million output tokens. The lowest qualifying quote in our catalog for this exact candidate is Vast.ai, Marketplace offer #41555529, Czechia, CZ, 3.11 USD/hour for the whole instance. It is listed on-demand hourly pricing with no minimum commitment; confirm stock and final charges. Stock status: listed. This is catalog coverage, not a whole-market comparison.

# Display values are rounded; totals use full stored rates.
Period budget ≈ 3.11 USD/hour × 720 hours × 1 replica(s) + 0.00 USD
              ≈ 2,235.84 USD
API USD/million output, including input = 0.4 + (4096 / 1024) × 0.2
                                       = 1.2
Spending investigation threshold = budget / API USD per million output × 1,000,000
                                ≈ 1.9 billion billed output tokens
Requests ≈ 1.8 million per 30 days
Model-specific traffic and cost calculation
Requests per period API bill, USD Candidate budget, USD Financial screen
100 thousand 122.88 2,235.84 Below candidate budget
1 million 1,228.80 2,235.84 Below candidate budget
1.5 million 1,843.20 2,235.84 Below candidate budget
2.2 million 2,703.36 2,235.84 Spend can fund a capacity test

At the spending threshold the workload demands 719 billed output tokens/second averaged over the calendar period, or 719 per replica if traffic is evenly shared. These are measurement targets, not measured GPU speeds. Replay the actual request lengths, bursts, reasoning settings and tool turns, then check aggregate output throughput, tail latency, failures and task success.

With the same 4,096 input tokens but 8,192 billed output tokens per request, the same budget corresponds to 550 thousand requests and 1,725 average output tokens/second. This longer-output sensitivity illustrates the effect of billed generation; it does not predict a reasoning setting's output length or the GPU's performance.

The period assumes 720 always-on hours and 1 replica(s). The explicit additional operating allowance is 0.00 USD. Include storage, network, taxes, monitoring, operational work and redundancy before comparing total service cost. When this allowance is zero the screen covers compute only. Extra replicas increase the bill; they are not a tested availability design.

API checked October 11, 2026 at 2:15 am EDT, collector source; rental checked October 11, 2026 at 1:01 pm EDT, quote source. Recorded terms: On-demand compute billed for actual active runtime; no monthly commitment. Storage persists and remains billable until deleted; transfer and other charges are separate. Quote notes: API discovery sample, capped at 12/GPU model. Rentable at check; capacity can change. Compute only; storage and traffic extra. Prices and calculations refresh automatically. Figures are rounded for readability; calculations use the full stored precision. Stale, retired or changed-identity observations suppress this calculation.

Self-hosting terms and evidence

The pinned Gemma 4 31B publisher model card declares Apache 2.0 and links to Google's Gemma 4 license page. Google's Apache 2.0 terms have been reviewed for this self-hosting comparison. When redistributing, provide the license, preserve applicable notices and identify modified files; follow any applicable NOTICE requirements. The license does not establish suitability or compliance for a particular application.

Publisher terms. Analysis reviewed October 10, 2026. Noach Ark did not rent this allocation or run inference, capacity or quality tests for this article. Prices and financial calculations refresh independently of this editorial review.