Run your model

Find a suitable rental

Choose your model and users, or specify technical requirements. Compare rental costs for your run duration.

01 Your requirements

Include weights, cached context and runtime overhead in your requirement. Memory is checked per GPU; multiple GPUs are not treated as pooled memory.

Count users making requests at the same time, with one active request per user. Idle users do not count. Results estimate memory needs; response speed is still unverified.

Conversation length and memory settings

Uses the sourced BF16 checkpoint and BF16 KV cache. Tokens include input and expected output for each user's request. The 4 GiB runtime reserve is an editable assumption; engine compatibility and actual memory use need verification.

Memory sizing: 8192 tokens per user · BF16 · 4 GiB runtime reserve.

02 Run duration and availability
Location preference · Any country

Set a location if your workload has latency or data-location requirements. We do not infer your required country.

Duration includes idle time while the instance is billable. Spot can stop unexpectedly; retries and lost work are additional.

03 Cost requirements

Budget applies to estimated whole-instance compute in the selected currency. Taxes and extras can increase the bill.

Storage and outbound transfer

For hourly rentals, blank storage retention uses your instance duration. For monthly rentals, enter retention hours if requesting storage. Monthly storage uses full billing months, defaulting to your monthly rental duration or one month for hourly rentals.

Reset requirements

Your workload, then your options

Compare options to run your model.

  1. Choose your model and parallel usersStart with one active user. Increase this to match simultaneous requests; the memory estimate includes each user's context.
  2. Set the time you will run itStart with 24 hours and no monthly commitment. Interruptible rentals are excluded unless you choose them.
  3. Compare estimated rental costsSee the whole instance cost for your duration, why each option matches and which charges still need checking.

Defaults are a starting scenario. Response speed, deployment compatibility and available capacity still need verification. How we handle pricing and evidence →