What to take away

  • Define the model, precision, context, concurrency, and service target before comparing GPUs.
  • Reject hardware that fails memory, software, host, or recovery requirements, regardless of its discount.
  • Buy used when verified capacity and performance meet the workload and the savings cover inspection, integration, and replacement risk.
  • Buy new when measured improvements or support reduce the cost of completing useful work enough to justify the premium.
  • Rent before buying when the workload is uncertain, occasional, or too large for an economical local system.

A GPU purchase is a decision about completing useful work within a budget and a deadline. The sticker price is only one part of that decision. A cheap card that cannot hold your model, support your software, or finish a job on time can become an expensive purchase.

Second-hand hardware can be an effective way to obtain memory capacity for experimentation, inference, and adapter fine-tuning. New hardware can justify its price through greater usable capacity, faster jobs, better energy efficiency, and support. These advantages must apply to your actual workload.

Use five gates before deciding how much a GPU is worth. A failure in any gate means rejecting the option or pricing a concrete remedy into the comparison.

Five qualification gates before purchasing a GPU
Decision gate Evidence to collect Reject or revise the option when
Workload fit Peak memory, model quality, latency, throughput, or training time The intended job cannot fit or meet its target
Software support Exact GPU, driver, framework, engine, and kernel compatibility The required stack is unsupported or impractical to maintain
Host compatibility Cooling, power, connectors, dimensions, PCIe, and system resources Integration exceeds the budget or cannot be made reliable
Reliability and recovery Condition, return rights, support, spare capacity, and restore procedure A failure would create an unacceptable service interruption
Total cost Installed cost, operating cost, useful work, recovery allowance, and resale assumptions A better qualified option costs less for the required outcome

Define the job before choosing the card. Write a workload specification that another engineer could reproduce. For inference, include the exact checkpoint and revision, quantization method, engine version, typical and maximum prompt lengths, output lengths, concurrency, and latency target. For training, add sequence length, batch size, optimizer, precision, checkpointing, and the number of parameters being updated.

An illustrative specification could be: an 8B model in a named 4-bit format, up to 8,192 tokens of total context, eight concurrent requests, and explicitly chosen limits for time to first token and token delivery. This is a benchmark specification, not a claim that every 8B model will work on a particular GPU. Test quality as well as speed: a quantized configuration that fails your evaluation is not a qualified substitute.

Inference and training workload requirements
Workload What usually drives the purchase What to measure before buying
Interactive inference Model memory, context, concurrency, and response latency Time to first token and token delivery under realistic requests
Production inference Throughput within latency targets, availability, and recovery Successful requests or tokens served within the target
Offline batch inference Completed work per hour and per unit of cost End-to-end batch completion time and output quality
Adapter fine-tuning Frozen weights, activations, sequence length, and trainable adapters Peak memory and time to an acceptable evaluation result
Full model training Weights, gradients, optimizer state, activations, and communication Stable step time, scaling, memory, and checkpoint recovery

Treat memory as a qualification gate. A useful first estimate is parameter count multiplied by bytes per stored parameter. The following numbers use decimal GB and describe raw dense weights only. They exclude metadata, runtime workspaces, activations, and the inference key-value cache.

Calculated raw dense model-weight sizes in decimal GB
Parameter count Raw FP16/BF16 weights, 2 bytes per parameter Raw packed 4-bit weights, 0.5 bytes per parameter
8 billion 16 GB 4 GB
14 billion 28 GB 7 GB
32 billion 64 GB 16 GB
70 billion 140 GB 35 GB

These are arithmetic estimates, not GPU fit recommendations. Actual quantization formats add overhead, and some components may remain at higher precision. Hugging Face distinguishes model size estimates from additional memory needed during execution. Compare measured usage and advertised capacity in consistent units; 1 GiB is approximately 1.074 decimal GB. Hugging Face model size estimator.

For inference, the key-value cache stores attention information from previous tokens. Longer sequences and more simultaneously active requests can increase its memory requirement. Capacity that appears sufficient for one short request may fail at the intended context and concurrency. Measure peak memory with representative requests and leave documented room for variation and runtime allocations. Transformers cache documentation.

Offloading some model or cache data to host memory may allow a workload to run, but it changes the performance comparison. Count transfer overhead and the additional RAM required. Test the resulting latency against your target before treating offloading as a solution. Transformers offloaded cache.

Full training requires a different calculation. One conventional mixed-precision AdamW configuration uses approximately 18 bytes per parameter for model weights, gradients, and optimizer state before activations and temporary buffers. That gives roughly 144 GB for an 8B model under those assumptions. Optimizer choice, precision, sharding, and offloading change the requirement. Transformers model memory anatomy.

Adapter fine-tuning updates a smaller set of parameters while keeping the base model frozen. QLoRA demonstrated fine-tuning a 65B model on a single 48 GB GPU in its published setup. That result supports a specific method and configuration; it does not guarantee that every 65B fine-tuning workload fits in 48 GB. QLoRA paper.

Compare useful performance. GPU specifications describe capabilities, but they do not establish the performance of your application. In LLM inference, processing the prompt and generating successive tokens can stress different resources. NVIDIA describes prompt processing as commonly compute intensive and token generation as often limited by memory bandwidth, with behavior affected by batching and configuration. NVIDIA inference optimization guide.

Use the same model revision, precision, engine, request distribution, and quality threshold when comparing candidates. Record power limits and host hardware. For interactive inference, compare latency at the intended concurrency. For a production endpoint, compare throughput while satisfying the latency target. For training, compare time to the required evaluation result rather than assuming a faster step produces the same model quality. vLLM provides benchmark tools for latency and serving performance. vLLM benchmark documentation.

Buy used when its capacity solves the job economically. Older hardware can remain useful when the workload fits, the software stack supports it, and its installed cost is substantially lower. As reference examples, NVIDIA lists the GeForce RTX 3090 with 24 GB of memory and 350 W graphics card power, and the RTX A6000 with 48 GB of ECC memory and a 300 W maximum power specification. These are manufacturer specifications, not measured workload consumption. RTX 3090 specifications, RTX A6000 specifications.

A qualified used 48 GB card may be a better purchase than a faster 32 GB card when the intended configuration needs more than 32 GB on one GPU. Capacity can avoid partitioning or offloading that would otherwise change the design. Conversely, spare VRAM has limited economic value if the job already fits comfortably and completion speed is the dominant requirement.

The savings must survive integration and acceptance testing. Require enough time to inspect and exercise the hardware, obtain a usable return agreement, and establish how a failure will be handled. An unexplained history is uncertainty to price into the deal; it is not proof that the card is defective.

Buy new when the improvement has a measurable value. New hardware can be justified by capacity, supported numerical formats, application performance, energy per job, or contractual support. NVIDIA lists the RTX 5090 with 32 GB of memory and 575 W total graphics power. The RTX PRO 6000 Blackwell family offers 96 GB ECC memory, with different cooling and power configurations across workstation, server, and Max-Q variants. Verify the exact product rather than purchasing from a family name. RTX 5090 specifications, RTX PRO 6000 family.

Do not infer energy efficiency from maximum power alone. A card that draws more power but finishes much sooner can use less energy per job. Measure average system power during the actual workload and multiply it by completion time. GPU-only measurements omit CPU, memory, storage, fans, and power-supply losses.

Hypothetical purchase and GPU electricity comparison
Illustrative comparison Used GPU New GPU
Purchase price $1,000 $1,800
Time for the same completed job 6 hours 4 hours
Assumed average GPU power during the job 350 W 250 W
GPU electricity at $0.20 per kWh $0.42 per job $0.20 per job

This hypothetical example saves $0.22 of GPU electricity per job. Electricity alone would need about 3,637 jobs to recover the $800 purchase premium. It excludes system power and all other costs. The two hours saved per job could be much more valuable if they unlock additional experiments or relieve a capacity bottleneck. Count that value only when the saved time changes an actual business outcome; unattended GPU time is not automatically paid engineer time.

Check the software before paying. An older GPU may have ample memory yet fail the requirements of a current inference engine or compiled kernel. NVIDIA's CUDA 13.0 release notes removed offline compilation and library support for Maxwell, Pascal, and Volta architectures; older toolchains may still be usable for a pinned legacy stack. vLLM's documented NVIDIA GPU installation requirement includes compute capability 7.5 or higher. These are different compatibility layers, so check every required component. CUDA 13.0 release notes, vLLM GPU installation.

For AMD hardware, verify the exact GPU, operating system, ROCm release, and framework against the supported configuration. Availability of a ROCm package does not establish support for every AMD card. Include engineering effort for unsupported configurations in the comparison. ROCm compatibility matrix.

Price the whole host. Confirm slot space, airflow, cooling design, power connectors, PSU capacity, CPU and RAM needs, PCIe layout, and storage. An accelerator designed for a server may require airflow a desktop cannot supply. For example, NVIDIA's A40 datasheet specifies passive cooling and an 8-pin CPU power connector. Validate the specified wiring rather than assuming a similarly shaped connector is interchangeable. NVIDIA A40 datasheet.

Two 24 GB GPUs do not automatically become one transparent 48 GB memory pool. Ordinary distributed data parallel training replicates the model across devices. Sharding methods can divide model state, but require appropriate software and introduce communication costs. Inference partitioning also needs engine support and an adequate interconnect. Benchmark the complete multi-GPU arrangement. PyTorch DistributedDataParallel, PyTorch Fully Sharded Data Parallel.

Check deployment terms as well as physical compatibility. NVIDIA's GeForce software license agreement states that GeForce and Titan software is not licensed for data center deployment. Verify the current agreement and applicable deployment rights for the planned environment during procurement. NVIDIA GeForce software license.

Give used hardware an acceptance procedure. Agree on the return deadline before purchase and complete these checks within it:

  1. Confirm the exact model, memory capacity, serial number, seller description, and sale terms. Keep the invoice and condition record.
  2. Inspect the board, connectors, fans, mounting hardware, and accessible surfaces for damage or modifications. Follow the seller's terms before disassembly.
  3. Install the card in a known working, correctly powered host with appropriate airflow and a supported driver.
  4. Exercise compute and memory with suitable diagnostics. Record temperature, clocks, power, throttling, resets, and available error counters.
  5. Run the intended inference or training workload repeatedly, including the target context, concurrency, or batch size. Confirm outputs and completed checkpoints.
  6. Save configuration, diagnostic, and workload logs so results can be compared and faults reproduced.
  7. Accept the card only after it meets the agreed requirements. Investigate failures promptly and use the return procedure if they remain unresolved.

Where supported, nvidia-smi exposes ECC and other device information. An unavailable field means unavailable information, not a zero error count. A failure can also originate in the driver, power supply, host memory, or cooling, so isolate the cause. Passing a stress test is evidence about behavior during the test; it does not establish remaining lifespan. NVIDIA System Management Interface documentation.

Do not assume warranty transfer. NVIDIA's own graphics-card warranty has purchaser and used-product exclusions; partner and seller policies can differ. Obtain the applicable terms for the specific card and transaction. NVIDIA graphics-card warranty.

For production, define recovery independently of the purchase label. A new GPU can fail, and a replacement warranty may take longer than the service can tolerate. Specify a tested spare, failover endpoint, rental contingency, or acceptable downtime. Include its cost. If recovery cannot meet the requirement, reject the deployment design.

Calculate ownership cost over a realistic period. Combine purchase price, installation and host upgrades, electricity and cooling, maintenance, support, and a documented recovery allowance, then subtract a conservative resale estimate. Do not treat forecast resale value as guaranteed cash. Compare the entire server when candidate GPUs require different hosts.

An additional 100 W sustained for a full year consumes 876 kWh, costing $175.20 at an assumed $0.20 per kWh. This is an arithmetic illustration, not a tariff or utilization forecast. For occasional jobs, energy per completed job is more useful than a continuous-operation estimate.

For used equipment, compare the discount with concrete allowances: inspection time, replacement fans or other permitted maintenance, host adaptation, return shipping, and the recovery plan. Avoid a universal percentage rule. A 40% discount can still be poor value if the card does not meet the workload; a smaller discount can be sensible if it supplies otherwise unaffordable usable capacity.

Rent when uncertainty or utilization makes ownership expensive. Rental capacity is useful for testing a candidate GPU, occasional training, workload exploration, and jobs that need a larger multi-GPU system. Compare completed work under matched quality and performance requirements. Include storage, transfers, startup, availability, and staff time where they differ.

Hypothetical rent-versus-own break-even assumptions
Illustrative rent-versus-own input Assumption
Net fixed ownership cost over 24 months $2,400, or $100 per month
Local variable cost $0.25 per active hour
Cloud cost for equal useful work per hour $2.00 per hour
Calculated monthly break-even $100 divided by ($2.00 minus $0.25), approximately 57 active hours

These hypothetical inputs assume equal useful output per hour and put relevant fixed costs into the $2,400 figure. They are not current market quotes. Different throughput invalidates an hour-for-hour comparison: use cost per completed job instead. Add omitted costs before making a purchase. Ownership becomes attractive above the illustrated utilization only if capacity, reliability, and the other gates are also satisfied.

A practical compromise is to own capacity for predictable baseline work and rent for bursts or large training runs. Confirm the rented environment can reproduce the required stack and that data transfers and availability do not erase the benefit.

Choose the purchase date from evidence. Buy when the workload is defined, qualified options are available, and delaying the purchase costs more than the expected benefit of waiting. Wait or rent when model choice, memory demand, or software support is still changing. Profile the current system first: a data-loading, CPU, storage, or network bottleneck may leave a GPU upgrade underused. Base launch-related timing on announced products, usable availability, and relevant measurements rather than speculative release dates.

Before approving a purchase, complete this worksheet. It turns a preference for new or used equipment into a reviewable decision.

GPU purchase approval worksheet
Approval question Evidence required
What exact job must it complete? Workload specification, model quality requirement, and service or completion target
Does it fit with operating room? Measured peak memory under representative and maximum intended conditions
Does it perform well enough? Reproducible workload results on the candidate or an equivalent configuration
Can we run and maintain the stack? GPU, driver, framework, engine, kernel, and operating-system compatibility
Can we install and use it correctly? Host, cooling, power, connector, interconnect, and deployment-rights checks
Is it the lowest qualified cost? Installed cost and operating estimate compared with new, used, and rented alternatives
What happens if it fails? Acceptance terms and a recovery plan that meets the downtime requirement

Choose used when verified condition, usable memory, performance, and support meet the job, and the installed-cost savings remain worthwhile after recovery and integration costs.

Choose new when measured capability or applicable support justifies the premium over qualified alternatives within the intended ownership period.

Choose rented capacity when the workload is uncertain, infrequent, temporarily large, or better served by a system you would struggle to keep productively occupied. The defensible decision is the option that completes the required work at an acceptable cost and risk.

Research date: October 4, 2026. Specifications and compatibility statements are sourced to vendor or project documentation. All purchase prices, power measurements, tariffs, job times, and rental rates in the cost examples are illustrative assumptions; no hardware acceptance, inference, training, or energy benchmark was run for this article. Recheck compatibility, licensing, warranty terms, and actual quotes before procurement.