What to take away
- Distinguish a monthly rental from an hourly bill, a monthly estimate, and a longer prepaid commitment.
- Compare the complete server configuration and total commitment—not just the GPU name or per-card price.
- Calculate the break-even point using billable rental hours, not GPU utilization.
- Include storage, transfer, taxes, setup costs, and cancellation terms in the comparison.
- Verify the actual offer and test your workload before committing to a month or longer.
Your model fits. The endpoint works. Now you want to leave it running for customers instead of launching it for an afternoon experiment.
An hourly GPU price suddenly becomes a recurring operating cost. A monthly rental looks attractive: one predictable payment and a machine available throughout the period. But a price followed by “/month” does not tell you what you are buying, how long you must commit, or whether it is cheaper than switching an hourly instance off when you do not need it.
Our GPU memory guide covers the first question: what configuration can support your workload? This guide tackles the next one: which rental arrangement makes sense once you have that shortlist?
Start with the billing model, not the price
Four offers can look similar on a pricing page and produce very different obligations:
- A monthly rental: you pay for a specified configuration for the rental period, whether you use it continuously or leave it idle. Check renewal and cancellation rules.
- An hourly service invoiced monthly: the invoice arrives monthly, but usage still determines the charge. The invoice schedule is not a discount.
- A monthly estimate: the provider displays what continuous hourly use would cost over an assumed number of hours. This is not necessarily a separate plan or a spending cap.
- A longer commitment expressed per month: the headline divides a multi-month price into smaller-looking units. The amount due and the commitment can be much larger.
GPU Mart illustrates the last distinction. Its site lists an RTX 4090 at $409/month, but its pricing note says displayed rates are monthly equivalents based on a 24-month billing term. A one-month option is shown; that does not establish that it costs $409. GPU Mart pricing
Runpod illustrates another category. Its current Pod documentation specifies three- or six-month prepaid savings plans. They cover GPU compute, exclude storage, and are non-refundable with fixed expiration dates. That may suit a steady service, but it is not a one-month trial. Runpod Pod pricing
Ask two separate questions: How much do I pay now, and what am I committed to paying overall?
Monthly options worth comparing
The following is a selected pricing snapshot checked on September 24, 2026, not a cheapest-provider ranking. Prices retain their original currencies; taxes, extras, stock, and final terms need confirmation. The configurations are not performance-equivalent.
| Provider | Advertised configuration | Advertised monthly price | Qualification |
|---|---|---|---|
| LeaderGPU | One RTX 4090, 24 GB | €329 | 64 GB system RAM and 480 GB SSD; VAT excluded. Availability must be checked. |
| HOSTKEY | One RTX 4090, 24 GB, gpu.v3-4090r | €350 | 16 vCPU, 64 GB RAM, 1 TB NVMe and 30 TB traffic; catalogue showed Buy now. |
| Latitude.sh | One H100, 80 GB, g3.h100.small | From $1,230 | 2 × 3.8 TB NVMe and 20 TB outbound transfer; confirm the region and billing selection. |
| Nexus Compute | RTX 4090 or RTX 5090 | From $250 per GPU | Dynamic reference prices; exact host configuration and bookable offer not verified. |
| GPU Mart | RTX 4090, 24 GB | $409 monthly equivalent | Headline rate assumes a 24-month term; one-month price not verified. |
These observations come from the providers' own pages: LeaderGPU, HOSTKEY, Latitude.sh, Nexus Compute, and GPU Mart.
The catalogue price is only the beginning. LeaderGPU's linked RTX 4090 order page showed the €329 configuration temporarily unavailable in the September 24 snapshot, with availability stated from September 25. That dated message was no longer shown when the source was rechecked September 25. The observed configuration included 10 TB of monthly transfer and listed €0.09/GB for additional traffic, excluding VAT. A dated availability message is not a capacity guarantee. Configuration and order details
HOSTKEY also listed a €279 RTX 4090 configuration, but marked it pre-order. That is not the same offer as the €350 configuration above. Its payment documentation describes monthly prepayment with hourly settlement on eligible cancellations; multi-month prepayments have different restrictions. Read the applicable terms before treating a prepaid month as either fully refundable or completely locked in. Catalogue, payment terms
XenCloud: private monthly quotes
XenCloud is known for exceptionally competitive monthly pricing. However, its offers are tailored privately and become available only after registration, so we cannot provide an official public price for this comparison. Request a quote directly from XenCloud to verify the configuration, terms, and final monthly rate. XenCloud
If a provider sends you a private offer, record the GPU, quantity, host resources, location, duration, total payment, and renewal rules. Compare the documented offer—not the provider's reputation or an estimate based on limited public information.
What live RTX 4090 marketplace prices looked like
At 08:11 EDT on September 25, 2026, we queried Vast.ai's official CLI for rentable, verified offers exposing one full 24 GB RTX 4090 with a reliability score above 95%. Four on-demand offers matched. With 5 GB of storage selected, they ranged from $0.457 to $0.549 per hour, with a median of $0.463. One interruptible offer matched, with a minimum bid of $0.667 per hour. At that moment, the available interruptible offer was more expensive than the cheapest on-demand offers. Vast.ai pricing model
Runpod's live public catalogue returned the same $0.34 hourly floor for a Community Cloud RTX 4090 on demand and for the minimum Spot bid. Its published Secure Cloud rate was $0.74 per hour. Spot capacity can be cheaper at other times, but this snapshot shows why the live bid and the on-demand alternative should be recorded together. Runpod pricing, Spot and on-demand instances
| Provider and mode | Observed price | 720-hour equivalent | Snapshot details |
|---|---|---|---|
| Vast.ai on demand | $0.457–$0.549/hour; $0.463 median | $329–$395; $334 median | Four verified, rentable, full-GPU offers above 95% reliability; 5 GB storage selected. |
| Vast.ai interruptible | $0.667/hour minimum bid | $480 | One matching bid offer; the instance may pause when outbid. |
| Runpod Community Cloud on demand | $0.34/hour floor | $245 | Lowest live public catalogue rate; availability can change. |
| Runpod Community Cloud Spot | $0.34/hour minimum bid | $245 | Interruptible; no discount versus the on-demand floor in this snapshot. |
| Runpod Secure Cloud on demand | $0.74/hour | $533 | Published RTX 4090 rate; storage is separate. |
The Vast.ai total includes the 5 GB storage selection used in the query; host-set transfer charges remain separate. Runpod storage is separate. The 720-hour figures are arithmetic equivalents, not fixed monthly offers or availability guarantees.
Public H100 prices across specialist and hyperscale clouds
The table below uses public prices checked on September 25, 2026. It separates the hourly rate from the minimum rentable configuration so a low per-GPU figure is not mistaken for a low invoice. The 720-hour column multiplies the observed hourly rate by 720; it does not convert Spot, Flex-start, or Capacity Block capacity into a monthly contract.
| Provider and offer | Public rate | 720-hour equivalent | Minimum configuration and qualification |
|---|---|---|---|
| Latitude.sh g3.h100.small | From $1,230/month | $1.71/hour equivalent | One H100 80 GB; 2 × 3.8 TB NVMe and 20 TB outbound transfer; region and billing selection need confirmation. |
| Runpod H100 PCIe Pod | $2.89/GPU-hour | $2,081 | One H100 80 GB, 16 vCPU and 188 GB RAM; storage is separate. |
| Lambda H100 PCIe | $3.29/GPU-hour | $2,369 | One H100 80 GB, 26 vCPU, 225 GiB RAM and 1 TiB SSD. |
| Nebius HGX H100 on demand | $3.85/GPU-hour | $2,772 | Published per-GPU rate with 16 vCPU and 200 GB RAM; confirm the deployable unit and region. |
| Nebius HGX H100 preemptible | From $0.79/GPU-hour | From $569 | Interruptible capacity; published per-GPU starting rate. |
| AWS p5.4xlarge Capacity Block | $5.191/hour | $3,738 | One H100 80 GB in selected regions; purchased as a scheduled Capacity Block. |
| Google Cloud a3-highgpu-1g Flex-start | $4.79/hour | $3,449 | One H100, 26 vCPU, 234 GiB RAM and 750 GiB local SSD; start time is flexible. |
| Google Cloud a3-highgpu-1g Spot | $6.620/hour | $4,767 | Same one-H100 configuration; interruptible and currently priced above Flex-start. |
| CoreWeave HGX H100 on demand | $49.24/node-hour; $6.155/GPU-hour | $35,453/node; $4,432/GPU | Eight-H100 node with 2,048 GB RAM and 61.44 TB local storage. |
| CoreWeave HGX H100 Spot | $19.51/node-hour; $2.439/GPU-hour | $14,047/node; $1,756/GPU | Eight-H100 interruptible node; the complete node is the minimum invoice. |
| Azure ND96isr H100 v5, East US | $98.32/node-hour; $12.29/GPU-hour | $70,790/node; $8,849/GPU | Eight-H100 node at Linux retail pricing from Azure's Retail Prices API. |
| Azure ND96isr H100 v5 Spot, East US | $18.17/node-hour; $2.271/GPU-hour | $13,082/node; $1,635/GPU | Eight-H100 interruptible node; no Spot SLA. |
Sources: Latitude.sh, Runpod, Lambda, Nebius, AWS Capacity Blocks, Google Cloud accelerator pricing, CoreWeave, and the Azure Retail Prices API.
These are not performance-equivalent machines. H100 PCIe and H100 SXM systems differ, and CPU, RAM, storage, networking, support, and capacity guarantees vary. The comparison is useful because it shows both the advertised rate and the amount of infrastructure you must actually rent.
Calculate the break-even hours
For the same usable configuration, with the same extras and service requirements, the simplest comparison is:
Break-even hours = monthly rental fee ÷ hourly rental rate.
Below that point, the hourly rental costs less. Above it, the monthly fee costs less. This assumes the hourly price stays fixed and that the resources can actually be released between sessions without undermining your service.
The table below applies that arithmetic to two pairs of Nexus Compute's published reference rates. They are illustrative inputs, not confirmed matching checkout offers. The provider explicitly says its rates are dynamic and per GPU. Nexus reference pricing
| Reference GPU | Monthly input | Hourly input | Calculated break-even | Share of a 720-hour month |
|---|---|---|---|---|
| RTX 4090 | $250 | $0.35 | 714.3 hours | 99.2% |
| H100 80 GB | $950 | $1.52 | 625.0 hours | 86.8% |
Calculations assume a 30-day month, unchanged rates, equivalent configurations, and no difference in additional charges. They do not predict throughput or confirm capacity.
The first example is instructive: at those inputs, 720 hours of hourly rental costs $252. The monthly fee saves just $2. At 100 hours, the hourly cost is $35, making the monthly commitment much harder to justify on rental cost alone.
Use the provider's actual billing period for your decision. A 30-day example is not a universal definition of a rental month.
Count rented hours, not busy hours
A GPU allocated to your service can spend time waiting for requests. If that allocated time remains billable under your plan, a low utilization reading does not reduce the rental charge.
For example, an endpoint kept allocated for all 720 hours of a 30-day month is a 720-hour rental even if its compute is busy for only part of the day. The relevant question is whether you can release capacity while still meeting your service requirements—not how often the GPU reaches a high utilization percentage.
For an internal batch job, waiting to restart may be acceptable. For a customer-facing endpoint, startup delays and uncertain replacement capacity may matter more. Runpod, for example, documents that stopping a Pod releases its GPU slot; that GPU may no longer be available when you restart. This is a provider-specific example, not a claim about every cloud. Runpod restart availability
A monthly commitment can therefore be valuable for operational reasons as well as price. Confirm what availability is actually guaranteed, including maintenance and hardware-failure handling; do not infer a service guarantee from the billing period.
Compare the complete configuration
The same GPU name does not make two rental offers identical. Record CPU allocation, system RAM, local and persistent storage, network speed, transfer allowance, and whether you receive a whole host or a virtual instance. For multi-GPU work, also verify how the devices are connected and whether your inference engine supports that arrangement.
Watch for per-card prices that require buying an entire node. LeaderGPU lists an eight-RTX-3090 configuration at €999/month, approximately €125 per GPU. Its single-RTX-3090 configuration is €249/month. The lower per-card figure is not a standalone €125 rental. LeaderGPU configuration table
Then test the same model, precision, input/output lengths, and concurrent workload on the candidates. A lower monthly bill is not a saving if the configuration cannot meet your response-time or capacity requirements. Keep rental-cost arithmetic separate from performance claims until you have measurements.
Add the costs outside the headline
Build two totals over the same period:
Monthly option = rental fee + setup costs allocated to that period + extras + applicable taxes.
Hourly option = hourly rate × billable hours + extras + applicable taxes.
Extras can include persistent storage, backups, outbound transfer, additional addresses, paid software, and support. Only include charges that actually apply to the quoted configuration; do not assume every provider bills each item separately.
For a one-month experiment, count the whole setup fee in that month. For a genuinely planned longer deployment, show both the upfront payment and the cost spread across the intended duration. Keep the same currency and tax treatment on both sides.
Cancellation is part of the comparison too. Check notice periods, automatic renewal, refund eligibility, and whether unused funds return to your payment method or remain as provider credit. If costs depend on usage tiers or a changing hourly rate, compare several realistic scenarios instead of relying on one break-even number.
A practical buying sequence
- Define the workload. Record the model, required memory, concurrency, response-time target, and hours when the service must be available.
- Get a complete offer. Capture the exact configuration, location, start date, billing period, amount due now, and total commitment. Save the quote date.
- Test before extending the term. Where a short rental is available, validate the configuration and deployment process before committing to a month or longer.
- Run the cost comparison. Include expected idle allocation, storage, transfer, taxes, and any migration or restart costs relevant to your operation.
- Plan the exit. Keep a backup outside the rented server and know how to cancel, retrieve your data, and move the service before renewal.
Monthly rental is most compelling when a suitable configuration needs to stay allocated for much of the period and the complete quote beats your realistic hourly alternative. It is less compelling when usage is occasional, the discount is small, or the workload is still changing.
Start with the workload, verify the offer, and calculate the commitment. The useful number is not the smallest price beside “/month.” It is the cost of keeping your actual service working for the period you need.
Method and sources: Public provider pages checked September 24, 2026; key pricing pages and the Vast.ai, Runpod, and Azure public pricing interfaces were rechecked September 25, 2026. Marketplace observations are timestamped snapshots using the filters stated above. Prices and availability are time-sensitive. No GPU rentals were purchased, no performance benchmarks were run, and provider reliability or capacity was not independently validated. Calculations are estimates using stated assumptions. Prepared with AI-assisted research and writing; the hero illustration is AI-generated. See Noach Ark's methodology.
