Cloud GPU for text embeddings
Select hardware around the embedding model, input length and batch size. The lowest hourly price matters only after that workload fits in memory.
The picks, with live prices
| Pick | GPU | VRAM | On-demand from | Where | |
|---|---|---|---|---|---|
| Price comparison | L4 | 24 GB | $0.44 | Jarvislabs on-demand | Rent → |
| Price comparison | RTX 4090 | 24 GB | $0.35 | TensorDock on-demand | Rent → |
| Price comparison | A100 | 80 GB | $0.89 | Jarvislabs on-demand | Rent → |
L4 Price comparison
Compare L4 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
RTX 4090 Price comparison
Compare RTX 4090 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
A100 Price comparison
Compare A100 only if its memory and software support meet your workload requirements. The table supplies the published capacity and current price; it does not claim this card runs your application faster.
Calculate compute cost per document
Enter the output and runtime from your own workload estimate or measurement. No throughput is assumed. Choose hardware only after checking your application requirements.
Compute per document = node hourly rate × runtime ÷ completed output. Updated rates 2026-10-02. Include startup, checkpoint and idle time in billed runtime. Storage, egress and other fees are additional.
Worth knowing
- Keep the exact embedding model and precision in the comparison; changing either can change vector outputs.
- Measure documents per second with your own input lengths before converting an hourly rate into a per-document cost.
- Separate initial indexing from query-time embedding. Their concurrency and uptime needs can differ.
FAQ
No. These are published hardware and rental-price comparisons. Supply your own memory, runtime and latency requirements before choosing a configuration.
Prices render from today's verified snapshot, not from when this guide was written. Full table on the homepage; break-even math in the calculator.