Runpod serverless pricing and other GPU clouds
Compare published per-second GPU rates with hourly instance prices. Idle workers, cold starts and minimum billing rules can change the bill; the label “serverless” does not establish those rules. Rates below show both units; the published unit and rounding differ by provider.
| GPU | $/hr equivalent | $/second | Provider | Updated | |
|---|---|---|---|---|---|
| Tesla T4 | $0.590 | $0.000164 | Modal | 2026-10-06 | Try → |
| Tesla T4 | $0.590 | $0.000164 | Cerebrium | 2026-10-06 | Try → |
| Tesla T4 | $0.631 | $0.000175 | Baseten | 2026-10-06 | Try → |
| L4 | $0.70 | $0.000194 | Koyeb | 2026-10-06 | Try → |
| RTX A6000 | $0.75 | $0.000208 | Koyeb | 2026-10-06 | Try → |
| L4 | $0.799 | $0.000222 | Modal | 2026-10-06 | Try → |
| L4 | $0.799 | $0.000222 | Cerebrium | 2026-10-06 | Try → |
| Tesla T4 | $0.81 | $0.000225 | Replicate | 2026-10-06 | Try → |
| L4 | $0.848 | $0.000236 | Baseten | 2026-10-06 | Try → |
| RTX 4090 | $1.10 | $0.000306 | RunPod | 2026-10-06 | Try → |
| A10 | $1.10 | $0.000306 | Cerebrium | 2026-10-06 | Try → |
| A10 | $1.10 | $0.000306 | Modal | 2026-10-06 | Try → |
| RTX PRO 4500 Blackwell | $1.15 | $0.000319 | RunPod | 2026-10-06 | Try → |
| L40S | $1.20 | $0.000333 | Koyeb | 2026-10-06 | Try → |
| A10 | $1.21 | $0.000335 | Baseten | 2026-10-06 | Try → |
| RTX 5090 | $1.58 | $0.000439 | RunPod | 2026-10-06 | Try → |
| A100 (unspecified) | $1.60 | $0.000444 | Koyeb | 2026-10-06 | Try → |
| L40S | $1.95 | $0.000542 | Modal | 2026-10-06 | Try → |
| L40S | $1.95 | $0.000542 | Cerebrium | 2026-10-06 | Try → |
| A100 PCIe 40GB | $2.00 | $0.000555 | Cerebrium | 2026-10-06 | Try → |
| A100 PCIe 40GB | $2.10 | $0.000583 | Modal | 2026-10-06 | Try → |
| A100 (unspecified) | $2.10 | $0.000583 | Cerebrium | 2026-10-06 | Try → |
| A100 SXM 80GB | $2.15 | $0.000597 | Koyeb | 2026-10-06 | Try → |
| RTX PRO 6000 (unspecified) | $2.20 | $0.000611 | Koyeb | 2026-10-06 | Try → |
| A100 (unspecified) | $2.50 | $0.000694 | Modal | 2026-10-06 | Try → |
| RTX PRO 6000 (unspecified) | $2.50 | $0.000694 | Cerebrium | 2026-10-06 | Try → |
| H100 (unspecified) | $2.50 | $0.000694 | Koyeb | 2026-10-06 | Try → |
| A100 (unspecified) | $2.72 | $0.000756 | RunPod | 2026-10-06 | Try → |
| H200 (SXM) | $3.00 | $0.000833 | Koyeb | 2026-10-06 | Try → |
| RTX PRO 6000 (unspecified) | $3.03 | $0.000842 | Modal | 2026-10-06 | Try → |
| H100 (unspecified) | $3.40 | $0.000944 | Cerebrium | 2026-10-06 | Try → |
| RTX PRO 6000 (unspecified) | $3.49 | $0.000969 | RunPod | 2026-10-06 | Try → |
| L40S | $3.51 | $0.000975 | Replicate | 2026-10-06 | Try → |
| H100 SXM | $3.95 | $0.001097 | Modal | 2026-10-06 | Try → |
| RTX PRO 6000 (unspecified) | $4.00 | $0.001111 | fal | 2026-10-06 | Try → |
| A100 (unspecified) | $4.00 | $0.001111 | Baseten | 2026-10-06 | Try → |
| H200 (SXM) | $4.20 | $0.001166 | Cerebrium | 2026-10-06 | Try → |
| H100 (unspecified) | $4.50 | $0.001250 | fal | 2026-10-06 | Try → |
| H200 (SXM) | $4.54 | $0.001261 | Modal | 2026-10-06 | Try → |
| H100 (unspecified) | $4.79 | $0.001331 | RunPod | 2026-10-06 | Try → |
| A100 (unspecified) | $5.04 | $0.001400 | Replicate | 2026-10-06 | Try → |
| H100 (unspecified) | $5.49 | $0.001525 | Replicate | 2026-10-06 | Try → |
| H200 (SXM) | $5.49 | $0.001525 | Replicate | 2026-10-06 | Try → |
| B200 | $5.50 | $0.001528 | Koyeb | 2026-10-06 | Try → |
| H200 (SXM) | $5.93 | $0.001647 | RunPod | 2026-10-06 | Try → |
| H200 (SXM) | $6.00 | $0.001667 | fal | 2026-10-06 | Try → |
| B200 | $6.01 | $0.001670 | Cerebrium | 2026-10-06 | Try → |
| B200 | $6.25 | $0.001736 | Modal | 2026-10-06 | Try → |
| H100 (unspecified) | $6.50 | $0.001806 | Baseten | 2026-10-06 | Try → |
| B300 | $7.10 | $0.001972 | Modal | 2026-10-06 | Try → |
| B200 | $7.99 | $0.002219 | fal | 2026-10-06 | Try → |
| H100 (unspecified) | $8.00 | $0.002222 | Fireworks AI | 2026-10-06 | Try → |
| H200 (SXM) | $8.00 | $0.002222 | Fireworks AI | 2026-10-06 | Try → |
| B200 | $8.64 | $0.002400 | RunPod | 2026-10-06 | Try → |
| B200 | $9.98 | $0.002772 | Baseten | 2026-10-06 | Try → |
| GB200 | $9.99 | $0.002775 | fal | 2026-10-06 | Try → |
| B300 | $12.99 | $0.003608 | fal | 2026-10-06 | Try → |
| B200 | $13.00 | $0.003611 | Fireworks AI | 2026-10-06 | Try → |
| B300 | $15.00 | $0.004167 | Fireworks AI | 2026-10-06 | Try → |
| GB300 | $20.00 | $0.005556 | Fireworks AI | 2026-10-06 | Try → |
Runpod Flex workers are billed from startup until fully stopped, rounded up to the nearest second. That includes loading the model, request execution and the idle timeout. Active workers stay running and require a sales inquiry for discounted rates. The table includes only individually named GPUs; mixed-card classes cannot promise a particular card.
Runpod publishes hourly equivalents on its pricing page; their per-second values here are those figures divided by 3,600 and may be rounded. Modal's per-second GPU rates are multiplied by 3,600 for comparison. Separately billed CPU, RAM and storage are excluded.
When serverless wins
Pick serverless when
- Traffic is bursty or unpredictable
- You serve a model occasionally (demos, side projects, internal tools)
- The provider’s scale-to-zero and billing rules match your idle periods
Pick an instance when
- Utilization is sustained (training, batch jobs, steady APIs)
- Cold starts are unacceptable
- You need full control of the environment or multi-GPU nodes
Modal includes a recurring monthly free credit tier; see free credits. RunPod also offers serverless endpoints priced separately from the pod rates shown on the RunPod page.
When does an always-on GPU cost less?
For the same GPU variant, divide the single-GPU instance rate by the serverless hourly equivalent. That fraction of an hour is the compute-only crossover. Request overhead and separately billed CPU, RAM, storage or warm workers can move it.
| GPU | Serverless $/second | Always-on $/hour | Crossover per hour |
|---|---|---|---|
| B200 | RunPod · $0.002400 | Lium · $5.40 | 37.5 active minutes (62.5% duty) |
| H200 (SXM) | RunPod · $0.001647 | DigitalOcean · $4.47 | 45.2 active minutes (75.4% duty) |
| RTX PRO 6000 (unspecified) | RunPod · $0.000969 | DataCrunch · $2.08 | 35.8 active minutes (59.7% duty) |
| H100 (unspecified) | RunPod · $0.001331 | Lium · $1.75 | 21.9 active minutes (36.5% duty) |
| A100 (unspecified) | RunPod · $0.000756 | Massed Compute · $1.35 | 29.8 active minutes (49.6% duty) |
| RTX PRO 4500 Blackwell | RunPod · $0.000319 | Massed Compute · $0.92 | 48.0 active minutes (80.0% duty) |
| B300 | fal · $0.003608 | Lium · $8.00 | 37.0 active minutes (61.6% duty) |
| B200 | fal · $0.002219 | Lium · $5.40 | 40.6 active minutes (67.6% duty) |
| H200 (SXM) | fal · $0.001667 | DigitalOcean · $4.47 | 44.7 active minutes (74.5% duty) |
| H100 (unspecified) | fal · $0.001250 | Lium · $1.75 | 23.3 active minutes (38.9% duty) |
| RTX PRO 6000 (unspecified) | fal · $0.001111 | DataCrunch · $2.08 | 31.3 active minutes (52.1% duty) |
| B300 | Modal · $0.001972 | Lium · $8.00 | No crossover within one hour; this serverless rate is lower even at full utilization |
| B200 | Modal · $0.001736 | Lium · $5.40 | 51.8 active minutes (86.4% duty) |
| H200 (SXM) | Modal · $0.001261 | DigitalOcean · $4.47 | 59.1 active minutes (98.5% duty) |
| H100 SXM | Modal · $0.001097 | Massed Compute · $3.14 | 47.7 active minutes (79.5% duty) |
| RTX PRO 6000 (unspecified) | Modal · $0.000842 | DataCrunch · $2.08 | 41.3 active minutes (68.8% duty) |
| A100 (unspecified) | Modal · $0.000694 | Massed Compute · $1.35 | 32.4 active minutes (54.0% duty) |
| A100 PCIe 40GB | Modal · $0.000583 | Lambda · $1.99 | 56.9 active minutes (94.8% duty) |
| L40S | Modal · $0.000542 | Massed Compute · $0.97 | 29.8 active minutes (49.7% duty) |
| A10 | Modal · $0.000306 | OVHcloud · $1.00 | 54.5 active minutes (90.8% duty) |
| L4 | Modal · $0.000222 | AWS · $0.805 | No crossover within one hour; this serverless rate is lower even at full utilization |
| A100 (unspecified) | Replicate · $0.001400 | Massed Compute · $1.35 | 16.1 active minutes (26.8% duty) |
| H100 (unspecified) | Replicate · $0.001525 | Lium · $1.75 | 19.1 active minutes (31.9% duty) |
| L40S | Replicate · $0.000975 | Massed Compute · $0.97 | 16.6 active minutes (27.6% duty) |
| H200 (SXM) | Replicate · $0.001525 | DigitalOcean · $4.47 | 48.9 active minutes (81.4% duty) |
| H100 (unspecified) | Fireworks AI · $0.002222 | Lium · $1.75 | 13.1 active minutes (21.9% duty) |
| H200 (SXM) | Fireworks AI · $0.002222 | DigitalOcean · $4.47 | 33.5 active minutes (55.9% duty) |
| B200 | Fireworks AI · $0.003611 | Lium · $5.40 | 24.9 active minutes (41.5% duty) |
| B300 | Fireworks AI · $0.004167 | Lium · $8.00 | 32.0 active minutes (53.3% duty) |
| GB300 | Fireworks AI · $0.005556 | DataCrunch · $10.32 | 31.0 active minutes (51.6% duty) |
| L4 | Baseten · $0.000236 | AWS · $0.805 | 56.9 active minutes (94.9% duty) |
| A10 | Baseten · $0.000335 | OVHcloud · $1.00 | 49.7 active minutes (82.8% duty) |
| A100 (unspecified) | Baseten · $0.001111 | Massed Compute · $1.35 | 20.3 active minutes (33.8% duty) |
| H100 (unspecified) | Baseten · $0.001806 | Lium · $1.75 | 16.2 active minutes (26.9% duty) |
| B200 | Baseten · $0.002772 | Lium · $5.40 | 32.5 active minutes (54.1% duty) |
| L4 | Cerebrium · $0.000222 | AWS · $0.805 | No crossover within one hour; this serverless rate is lower even at full utilization |
| A10 | Cerebrium · $0.000306 | OVHcloud · $1.00 | 54.5 active minutes (90.8% duty) |
| A100 PCIe 40GB | Cerebrium · $0.000555 | Lambda · $1.99 | 59.8 active minutes (99.6% duty) |
| L40S | Cerebrium · $0.000542 | Massed Compute · $0.97 | 29.8 active minutes (49.7% duty) |
| A100 (unspecified) | Cerebrium · $0.000583 | Massed Compute · $1.35 | 38.6 active minutes (64.3% duty) |
| H100 (unspecified) | Cerebrium · $0.000944 | Lium · $1.75 | 30.9 active minutes (51.5% duty) |
| H200 (SXM) | Cerebrium · $0.001166 | DigitalOcean · $4.47 | No crossover within one hour; this serverless rate is lower even at full utilization |
| B200 | Cerebrium · $0.001670 | Lium · $5.40 | 53.9 active minutes (89.8% duty) |
| RTX PRO 6000 (unspecified) | Cerebrium · $0.000694 | DataCrunch · $2.08 | 50.1 active minutes (83.5% duty) |
| L4 | Koyeb · $0.000194 | AWS · $0.805 | No crossover within one hour; this serverless rate is lower even at full utilization |
| RTX A6000 | Koyeb · $0.000208 | Massed Compute · $0.57 | 45.6 active minutes (76.0% duty) |
| L40S | Koyeb · $0.000333 | Massed Compute · $0.97 | 48.5 active minutes (80.8% duty) |
| A100 (unspecified) | Koyeb · $0.000444 | Massed Compute · $1.35 | 50.6 active minutes (84.4% duty) |
| A100 SXM 80GB | Koyeb · $0.000597 | Massed Compute · $1.38 | 38.5 active minutes (64.2% duty) |
| RTX PRO 6000 (unspecified) | Koyeb · $0.000611 | DataCrunch · $2.08 | 56.9 active minutes (94.8% duty) |
| H100 (unspecified) | Koyeb · $0.000694 | Lium · $1.75 | 42.0 active minutes (70.0% duty) |
| H200 (SXM) | Koyeb · $0.000833 | DigitalOcean · $4.47 | No crossover within one hour; this serverless rate is lower even at full utilization |
| B200 | Koyeb · $0.001528 | Lium · $5.40 | 58.9 active minutes (98.2% duty) |
At a request duration you supply, requests at crossover = active seconds ÷ billed seconds per request. This assumes no overlap between requests and no idle worker charges. Per-second GPU pricing does not by itself guarantee zero idle charges.