Serverless GPU pricing
Serverless GPU = per-second billing that scales to zero between requests: you pay for compute seconds, not idle hours. For bursty inference it beats renting an instance; for sustained load an on-demand instance is usually cheaper. Rates below are $/hr equivalents of per-second prices.
| GPU | $/hr equivalent | $/second | Provider | Verified | |
|---|---|---|---|---|---|
| Tesla T4 | $0.59 | $0.000164 | Modal | 2026-08-22 | Try → |
| Tesla T4 | $0.59 | $0.000164 | Cerebrium | 2026-08-22 | Try → |
| Tesla T4 | $0.63 | $0.000175 | Baseten | 2026-08-22 | Try → |
| L4 | $0.70 | $0.000194 | Koyeb | 2026-08-22 | Try → |
| RTX A6000 | $0.75 | $0.000208 | Koyeb | 2026-08-22 | Try → |
| L4 | $0.80 | $0.000222 | Modal | 2026-08-22 | Try → |
| L4 | $0.80 | $0.000222 | Cerebrium | 2026-08-22 | Try → |
| Tesla T4 | $0.81 | $0.000225 | Replicate | 2026-08-22 | Try → |
| L4 | $0.85 | $0.000236 | Baseten | 2026-08-22 | Try → |
| A10 | $1.10 | $0.000306 | Cerebrium | 2026-08-22 | Try → |
| A10 | $1.10 | $0.000306 | Modal | 2026-08-22 | Try → |
| L40S | $1.20 | $0.000333 | Koyeb | 2026-08-22 | Try → |
| A10 | $1.21 | $0.000335 | Baseten | 2026-08-22 | Try → |
| A100 (unspecified) | $1.60 | $0.000444 | Koyeb | 2026-08-22 | Try → |
| L40S | $1.95 | $0.000542 | Modal | 2026-08-22 | Try → |
| L40S | $1.95 | $0.000542 | Cerebrium | 2026-08-22 | Try → |
| A100 PCIe 40GB | $2.00 | $0.000555 | Cerebrium | 2026-08-22 | Try → |
| A100 PCIe 40GB | $2.10 | $0.000583 | Modal | 2026-08-22 | Try → |
| A100 (unspecified) | $2.10 | $0.000583 | Cerebrium | 2026-08-22 | Try → |
| A100 SXM 80GB | $2.15 | $0.000597 | Koyeb | 2026-08-22 | Try → |
| RTX PRO 6000 (unspecified) | $2.20 | $0.000611 | Koyeb | 2026-08-22 | Try → |
| A100 (unspecified) | $2.50 | $0.000694 | Modal | 2026-08-22 | Try → |
| RTX PRO 6000 (unspecified) | $2.50 | $0.000694 | Cerebrium | 2026-08-22 | Try → |
| H100 (unspecified) | $2.50 | $0.000694 | Koyeb | 2026-08-22 | Try → |
| RTX PRO 6000 (unspecified) | $2.99 | $0.000831 | fal | 2026-08-22 | Try → |
| H200 (SXM) | $3.00 | $0.000833 | Koyeb | 2026-08-22 | Try → |
| RTX PRO 6000 (unspecified) | $3.03 | $0.000842 | Modal | 2026-08-22 | Try → |
| H100 (unspecified) | $3.40 | $0.000944 | Cerebrium | 2026-08-22 | Try → |
| L40S | $3.51 | $0.000975 | Replicate | 2026-08-22 | Try → |
| H100 SXM | $3.95 | $0.001097 | Modal | 2026-08-22 | Try → |
| A100 (unspecified) | $4.00 | $0.001111 | Baseten | 2026-08-22 | Try → |
| H200 (SXM) | $4.20 | $0.001166 | Cerebrium | 2026-08-22 | Try → |
| H200 (SXM) | $4.50 | $0.001250 | fal | 2026-08-22 | Try → |
| H100 (unspecified) | $4.50 | $0.001250 | fal | 2026-08-22 | Try → |
| H200 (SXM) | $4.54 | $0.001261 | Modal | 2026-08-22 | Try → |
| A100 (unspecified) | $5.04 | $0.001400 | Replicate | 2026-08-22 | Try → |
| H100 (unspecified) | $5.49 | $0.001525 | Replicate | 2026-08-22 | Try → |
| H200 (SXM) | $5.49 | $0.001525 | Replicate | 2026-08-22 | Try → |
| B200 | $5.50 | $0.001528 | Koyeb | 2026-08-22 | Try → |
| B200 | $6.01 | $0.001670 | Cerebrium | 2026-08-22 | Try → |
| B200 | $6.25 | $0.001736 | Modal | 2026-08-22 | Try → |
| B200 | $6.25 | $0.001736 | fal | 2026-08-22 | Try → |
| H100 (unspecified) | $6.50 | $0.001806 | Baseten | 2026-08-22 | Try → |
| H100 (unspecified) | $7.00 | $0.001944 | Fireworks AI | 2026-08-22 | Try → |
| H200 (SXM) | $7.00 | $0.001944 | Fireworks AI | 2026-08-22 | Try → |
| B300 | $7.10 | $0.001972 | Modal | 2026-08-22 | Try → |
| B300 | $8.50 | $0.002361 | fal | 2026-08-22 | Try → |
| B200 | $9.98 | $0.002772 | Baseten | 2026-08-22 | Try → |
| B200 | $10.00 | $0.002778 | Fireworks AI | 2026-08-22 | Try → |
| B300 | $12.00 | $0.003333 | Fireworks AI | 2026-08-22 | Try → |
| GB300 | $18.00 | $0.005000 | Fireworks AI | 2026-08-22 | Try → |
Serverless bills only active compute seconds but adds cold-start latency and per-request overhead. Sustained workloads above roughly 50-60% utilization are cheaper on-demand: run your numbers in the calculator.
When serverless wins
Pick serverless when
- Traffic is bursty or unpredictable
- You serve a model occasionally (demos, side projects, internal tools)
- You want zero idle cost and no capacity management
Pick an instance when
- Utilization is sustained (training, batch jobs, steady APIs)
- Cold starts are unacceptable
- You need full control of the environment or multi-GPU nodes
Modal includes a recurring monthly free credit tier; see free credits. RunPod also offers serverless endpoints priced separately from the pod rates shown on the RunPod page.