Cloud GPU for batch inference
Choose the lowest total compute cost among configurations that meet your completion deadline and model memory requirement. An hourly price is only one input.
The picks, with live prices
| Pick | GPU | VRAM | On-demand from | Where | |
|---|---|---|---|---|---|
| Price comparison | RTX 4090 | 24 GB | $0.35 | TensorDock on-demand | Rent → |
| Price comparison | L40S | 48 GB | $0.97 | Massed Compute on-demand | Rent → |
| Price comparison | A100 | 80 GB | $0.89 | Jarvislabs on-demand | Rent → |
| Price comparison | H100 | 94 GB | $1.99 | Voltage Park on-demand | Rent → |
RTX 4090 Price comparison
Compare RTX 4090 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
L40S Price comparison
Compare L40S only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
A100 Price comparison
Compare A100 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
H100 Price comparison
Compare H100 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
Calculate compute cost per completed item
Enter the output and runtime from your own workload estimate or measurement. No throughput is assumed. Choose hardware only after checking your application requirements.
Compute per completed item = node hourly rate × runtime ÷ completed output. Updated rates 2026-10-02. Include startup, checkpoint and idle time in billed runtime. Storage, egress and other fees are additional.
Worth knowing
- Estimate runtime from a representative batch you measure; no tokens-per-second benchmark is assumed here.
- Include data-loading, checkpoint and output-write time in billed hours.
- Use interruptible rates only when the job can resume and your deadline allows your own restart allowance.
FAQ
No. These are published hardware and rental-price comparisons. Supply your own memory, runtime and latency requirements before choosing a configuration.
Prices render from today's verified snapshot, not from when this guide was written. Full table on the homepage; break-even math in the calculator.