Data as of: September 28, 2026. Prices are public listed compute rates, not verified checkout quotes or available capacity.
What to take away
- Measure the spot discount against the same provider's on-demand price for the same configuration and region. The selected offers below range from an 82% discount to no discount.
- Check the provider's reclaim process and monitor its signals. Notice and shutdown windows differ, and they do not guarantee time to recover.
- Include billed startup time in the cost comparison. Under the stated assumptions, spot saves money only when its share of time ready to serve exceeds the spot-to-on-demand price ratio.
- Measure both cold-start time and the wait for replacement capacity. A low effective compute rate can coexist with an unacceptable service gap.
- Keep enough running on-demand capacity for the traffic you must serve when every spot replica is unavailable.
Your endpoint works, the model fits, and you know roughly what a million output tokens costs on an on-demand rental. Then you notice the spot column: the same eight-GPU node for less than half the hourly price.
The discount buys a different service. The provider can reclaim spot capacity. For a live endpoint, that can mean interrupted streams, lost work, and a wait for another node to load the model.
Our monthly GPU rentals guide covers paying for capacity over a fixed period. This guide covers paying less for capacity that can disappear.
Spot, preemptible, interruptible: read the rules
Providers use different names for reclaimable capacity. Before comparing prices, establish four things:
- How the price is set: a published rate, a bid, or another mechanism. For example, Vast.ai uses bids for interruptible instances. Vast.ai pricing
- How interruption is signaled: the event or metadata your service must monitor, the notice window, and the shutdown deadline.
- What remains afterward: a stopped VM, retained disks, or deleted resources. Retained storage may continue to incur charges.
- How replacement happens: automatic replenishment, a restart you initiate, or a new allocation request that may have to wait.
Do not assume that an evicted instance returns automatically. Azure says it does not; CoreWeave keeps a fully reclaimed Spot Node Pool active and attempts to replenish it when capacity returns. Azure Spot VMs, CoreWeave Spot Node Pools
How big is the discount?
Compare spot and on-demand rates from the same provider for the same SKU, region, GPU count, and host configuration, checked together. Otherwise, you may be measuring differences in hardware or packaging instead of a spot discount.
The following selected rates were checked on September 28, 2026. All prices are USD per complete eight-GPU instance-hour, rounded to cents. Storage, transfer, taxes, and other separately billed services are excluded.
| Provider and instance | Region | On-demand | Spot | Spot discount |
|---|---|---|---|---|
| Azure ND96isr H100 v5, 8 × H100 | East US | $98.32 | $18.17 | 82% |
| Azure ND96isr MI300X v5, 8 × MI300X | Central US | $59.04 | $10.91 | 82% |
| CoreWeave HGX H100, 8 GPUs | North America | $49.24 | $19.71 | 60% |
| CoreWeave HGX H200, 8 GPUs | North America | $50.44 | $20.93 | 59% |
| CoreWeave HGX B200, 8 GPUs | North America | $68.80 | $34.11 | 50% |
| Google Cloud a3-ultragpu-8g, 8 × H200 | Iowa (us-central1) | $84.81 | $50.87 | 40% |
| Azure ND96isr H200 v5, 8 × H200 | East US 2 | $84.80 | $84.80 | 0% |
Spot discount = 1 − spot rate ÷ on-demand rate. Percentages use the source rates before rounding.
Azure prices come from separate Linux Consumption meters in the Retail Prices API: H100 in East US, MI300X in Central US, and H200 in East US 2. Other rows come from CoreWeave's North American rate table and Google Cloud's accelerator-optimized pricing table, with Iowa selected and hourly billing displayed.
Two things stand out. First, the discount belongs to an offer, not to a provider: the selected Azure H100 and MI300X offers save 82%, while its H200 offer saves nothing.
Second, the deepest discount does not necessarily identify the lowest spot rate. Azure's H100 rate is lower than CoreWeave's despite its higher on-demand baseline. Compare both the discount within each offer and the resulting rates across your shortlist. Matching GPU names alone do not establish equal throughput, networking, or service terms.
At equal spot and on-demand rates, reclaim risk buys no compute saving. More broadly, a marketplace's cheapest on-demand listing and cheapest interruptible listing may come from different hosts. Treat that as a comparison of alternatives, not a matched spot discount. Use the rental comparison to inspect each offer's configuration and source check.
Price the interruptions, not just the rate
An interruption creates two separate costs.
Billed recovery time: a replacement may bill while it pulls the serving image, reads model files, loads weights into GPU memory, and warms up. That is paid capacity not yet ready to serve.
Missing capacity: you may wait for an allocation before billing starts. The compute bill can fall while your endpoint's available capacity falls too. The rate calculation below does not price delayed requests, lost revenue, fallback service, or engineering effort.
Start with a simple capacity-cost model. Let S be the spot rate per node-hour, B the billed node-hours, and U the node-hours ready to serve the chosen workload:
Effective spot compute rate per ready node-hour = S × B ÷ U.
Ready time includes time waiting for requests while the model is loaded and healthy. It is not the GPU's busy-time percentage. Keep that definition consistent for both rental options.
If the equivalent on-demand node costs O per hour and is ready for all its billed hours, spot has a lower cost per ready node-hour when:
U ÷ B > S ÷ O.
This assumes equal usable throughput when ready, a constant rate, and no difference in extras. If on-demand also spends billed time starting or recovering, compare O × B_on-demand ÷ U_on-demand with the effective spot rate. If rates change, use the actual compute charge divided by ready node-hours.
For the selected offers, the break-even ready share is approximately 0.18 for Azure H100, 0.40 for CoreWeave H100, and 0.60 for Google Cloud H200. Azure H200 requires a share greater than 1.00, which is impossible under this model: even perfect readiness only ties its on-demand rate.
For the customer-facing economic result, measure total cost per successfully delivered token at your latency target. That captures retry work, demand, and throughput changes that ready hours alone cannot describe. See our cost per million tokens guide.
Worked example: an eight-H100 node
Use the CoreWeave H100 rates above: $19.71/hour spot and $49.24/hour on-demand. Suppose each spot replica is reclaimed three times per day and each replacement spends 20 billed minutes starting up.
For this illustration, assume replacement capacity is available immediately, there is no billing overlap, and the replica accrues 24 billed hours in the day. These are assumptions, not observed interruption or recovery measurements.
- Billed time: 24 node-hours.
- Ready time: 24 − (3 × 20 ÷ 60) = 23 node-hours.
- Effective spot rate: $19.71 × 24 ÷ 23 = $20.57 per ready node-hour.
- Saving against a fully ready on-demand node: approximately 58%.
The discount survives the assumed startup overhead. It does not establish that a service running on that replica alone would be acceptable: it still loses an hour of ready capacity, plus any unmodeled wait for replacement.
Reload time is a workload variable
Model size can make recovery slow. Meta's pinned BF16 checkpoint metadata gives approximately 141.1 GB of weight data for Llama 3.3 70B Instruct and 811.7 GB for Llama 3.1 405B Instruct. These are decimal GB, derived from the BF16 parameter counts.
At an illustrative aggregate read rate of 1 GB/s, reading that much weight data takes about 2.4 minutes and 13.5 minutes, respectively. These are file-read estimates, not cold-start measurements. Provisioning, image pulls, GPU loading, and warm-up add work; caching, quantization, and parallel reads can change the result.
The 405B BF16 example is not a deployment recommendation for the eight-H100 node: its weights alone exceed that node's nominal 640 GB of GPU memory. Use a configuration that fits the actual checkpoint and runtime overhead, as explained in our GPU memory guide.
Keep a running on-demand floor
For a live endpoint, a mixed fleet lets on-demand replicas cover your minimum service level while spot replicas add capacity above it. Size that floor for the traffic you must serve after all spot replicas disappear, including any retry load.
Using the worked example, each on-demand node supplies 24 ready hours per day and each spot replica supplies 23. The fleet's effective rate is total compute cost ÷ total ready node-hours, so all nodes use the same denominator:
| Fleet | Effective compute cost per ready node-hour | Saving vs a fully ready on-demand fleet | On-demand nodes remaining if all spot nodes are reclaimed |
|---|---|---|---|
| 2 on-demand | $49.24 | — | 2 |
| 1 on-demand + 1 spot | $35.21 | 28% | 1 |
| 1 on-demand + 3 spot | $27.97 | 43% | 1 |
For example, the two-node mixed fleet costs ($49.24 + $19.71) × 24 = $1,654.80/day and supplies 24 + 23 = 47 ready node-hours: $1,654.80 ÷ 47 = $35.21. The four-node fleet costs $2,600.88/day and supplies 93 ready node-hours. Values are rounded after calculation.
These are hypothetical capacity-cost averages, not equal-throughput or uptime guarantees. Each saving compares with fully ready on-demand capacity using the same cost metric. A running on-demand floor avoids spot eviction risk; ordinary failures and maintenance still require redundancy. A floor of one node only preserves the traffic one healthy node can serve.
Design the service to survive a reclaim
Start with the provider's actual sequence. Documentation checked September 28, 2026 describes the following:
| Provider | Notice and shutdown process | Reclaim behavior |
|---|---|---|
| AWS EC2 Spot | Best-effort two-minute notice for stopping or termination. Hibernation, where supported, begins immediately. | Stop, terminate, or hibernate, as configured. |
| Google Cloud Spot VMs | Default: no dedicated delay before shutdown. Optional 120-second preemption notice is in Preview, followed by a best-effort shutdown period of up to 30 seconds. | Stop by default, or delete if selected. Retained persistent disks remain billable. |
| Azure Spot VMs | Opt-in in-VM Scheduled Events notifications, delivered best effort up to 30 seconds before eviction. | Deallocate by default, retaining billable disks, or delete. No automatic restart. |
| CoreWeave CKS Spot Node Pools | Warning and cordon first; drain starts after two minutes. Pod shutdown uses its configured grace period, subject to a maximum seven-minute preemption window. | Node leaves when drained or the window expires. The pool attempts to replenish capacity when available. |
Sources: AWS interruption notices, Google Cloud preemption process, Azure eviction policy, and CoreWeave preemption timeline.
CoreWeave's seven-minute maximum is not seven minutes of unrestricted serving: draining begins after two minutes, and a Pod can terminate earlier. Kubernetes cordoning prevents new Pods from being scheduled; your application must also stop routing new inference requests to the replica.
Treat notice as a chance to drain requests and begin recovery. Only rely on a replacement becoming ready before shutdown if you have tested that complete path and have capacity available. Also handle abrupt loss without a notice.
- Detect and drain. Monitor the provider signal, remove the replica from request routing, and let requests finish within the remaining shutdown budget. AWS recommends checking its interruption notice every five seconds. AWS polling guidance
- Keep durable state outside serving nodes. Losing a replica loses its in-memory KV cache and active streams. Retain request context where another replica can access it.
- Make retries explicit. Before any output is delivered, the gateway may retry within a bounded deadline. After a partial stream, report the interruption and support a deliberate restart; do not silently replay output or assume identical text. Deduplicate requests that can trigger external side effects, and cap retries so they do not overwhelm the floor.
- Reduce and measure cold starts. Pin the checkpoint and serving image, keep weights near the GPUs, and measure allocation-to-readiness time. Separately measure the wait for allocation. Use the billed portion in the rate calculation.
- Control retained-resource charges. Choose retention or deletion deliberately and check disk deletion settings. Persistent storage and backups need their own cost line.
- Test alternative capacity pools. Where your workload permits, qualify other zones or instance types. Benchmark each alternative before relying on it, and plan for several spot replicas to be reclaimed together.
- Protect the floor. Keep the minimum capacity running and ready. Set priorities, admission limits, or load shedding for traffic beyond what it can serve during recovery.
A practical sequence
- Record the matched rates. Capture configuration, region, currency, billing basis, extras, and check time.
- Calculate spot ÷ on-demand. At or above 1.00, there is no headline compute saving. Near 1.00, modest recovery overhead can erase it.
- Measure your cold start. Record billed startup time and total request-to-readiness time separately, using the actual model, image, and storage.
- Run a trial. Count reclaims, unavailable-capacity periods, and recovery durations. Several days can provide an initial sample; a quiet trial does not establish a future interruption rate.
- Compare effective and total costs. Calculate actual compute charges per ready node-hour, then evaluate delivered-token costs, extras, fallback capacity, and operational effort.
- Size the floor and rehearse failure. Verify that essential traffic meets your latency target after all spot replicas are removed. Exercise the provider's eviction signal as well as abrupt node loss, and check clients, routing, retries, and replacement logic. Our GPU benchmarking guide explains how to measure service behavior.
- Recheck prices and behavior. Rates, capacity, and product rules can change. Revisit the decision when the discount or measured recovery cost changes.
Spot is worth considering when the rate reduction survives recovery overhead and the remaining fleet can serve the traffic you must keep. The decision rests on both cost and service behavior: a cheaper ready hour cannot compensate for an endpoint that misses its requirements.
Method and limitations
The seven price rows are selected public rate-card observations checked on September 28, 2026, not a complete market survey or verified rental quotes. Azure's Spot meters were kept separate from Low Priority meters. Google's row uses the displayed Iowa hourly configuration; CoreWeave's rows use its North American table. Primary sources are linked beside the comparisons and provider rules.
Discounts and capacity-cost figures are calculated estimates. The worked example assumes three interruptions per day, 20 billed startup minutes per interruption, immediate replacement capacity, no overlapping billing, and equal throughput when ready. Model-read estimates assume an aggregate 1 GB/s and BF16 weight data. No GPU rentals were purchased and no cold starts, interruptions, or inference performance were benchmarked for this article. See Noach Ark's methodology.
Production note: Prepared with AI-assisted research and editing. Prices are dated public observations. Calculations are illustrative estimates. The hero is an AI-generated conceptual illustration, not a hardware schematic.
