Cloud GPU for self-hosted AI agents
Price the model-serving workload separately from the agent’s tools and orchestration. Renting a GPU is useful only when you need to host a model rather than call an API.
The picks, with live prices
| Pick | GPU | VRAM | On-demand from | Where | |
|---|---|---|---|---|---|
| Price comparison | L4 | 24 GB | $0.44 | Jarvislabs on-demand | Rent → |
| Price comparison | L40S | 48 GB | $0.97 | Massed Compute on-demand | Rent → |
| Price comparison | A100 | 80 GB | $0.89 | Jarvislabs on-demand | Rent → |
| Price comparison | H100 | 94 GB | $1.99 | Voltage Park on-demand | Rent → |
L4 Price comparison
Compare L4 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
L40S Price comparison
Compare L40S only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
A100 Price comparison
Compare A100 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
H100 Price comparison
Compare H100 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
Calculate compute cost per completed task
Enter the output and runtime from your own workload estimate or measurement. No throughput is assumed. Choose hardware only after checking your application requirements.
Compute per completed task = node hourly rate × runtime ÷ completed output. Updated rates 2026-10-02. Include startup, checkpoint and idle time in billed runtime. Storage, egress and other fees are additional.
Worth knowing
- Use the model memory calculator for the exact model and context length.
- Budget GPU uptime as well as active token generation; a process waiting on an external tool may leave a rented GPU idle.
- Compare the expected token bill with the full rental bill before choosing self-hosting.
FAQ
No. These are published hardware and rental-price comparisons. Supply your own memory, runtime and latency requirements before choosing a configuration.
Prices render from today's verified snapshot, not from when this guide was written. Full table on the homepage; break-even math in the calculator.