LLM VRAM requirements
Model weights are only part of GPU memory use. Compare weight sizes below, then open a model to account for context, runtime overhead and training memory.
| Model | Parameters | 4-bit weights, GB | Cache estimate |
|---|---|---|---|
| Qwen2.5 0.5B | 0.49 billion | 0.2 | Config supported |
| Qwen2.5 Coder 0.5B | 0.49 billion | 0.2 | Config supported |
| Qwen3 0.6B | 0.75 billion | 0.4 | Config supported |
| gemma 3 1b | 1.00 billion | 0.5 | Not available |
| Llama 3.2 1B | 1.24 billion | 0.6 | Not available |
| Qwen2.5 1.5B | 1.54 billion | 0.8 | Config supported |
| Qwen2.5 Coder 1.5B | 1.54 billion | 0.8 | Config supported |
| DeepSeek R1 Distill Qwen 1.5B | 1.78 billion | 0.9 | Config supported |
| Qwen3 1.7B | 2.03 billion | 1.0 | Config supported |
| gemma 2 2b | 2.61 billion | 1.3 | Not available |
| Qwen2.5 3B | 3.09 billion | 1.5 | Config supported |
| Qwen2.5 Coder 3B | 3.09 billion | 1.5 | Config supported |
| Llama 3.2 3B | 3.21 billion | 1.6 | Not available |
| Phi 3 mini 4k | 3.82 billion | 1.9 | Not available |
| Phi 4 mini | 3.84 billion | 1.9 | Not available |
| Qwen3 4B | 4.02 billion | 2.0 | Config supported |
| gemma 3 4b | 4.30 billion | 2.2 | Not available |
| Mistral 7B v0.3 | 7.25 billion | 3.6 | Config supported |
| Qwen2.5 7B | 7.62 billion | 3.8 | Config supported |
| Qwen2.5 Coder 7B | 7.62 billion | 3.8 | Config supported |
| DeepSeek R1 Distill Qwen 7B | 7.62 billion | 3.8 | Config supported |
| Llama 3.1 8B | 8.03 billion | 4.0 | Not available |
| DeepSeek R1 Distill Llama 8B | 8.03 billion | 4.0 | Config supported |
| Qwen3 8B | 8.19 billion | 4.1 | Config supported |
| gemma 2 9b | 9.24 billion | 4.6 | Not available |
| gemma 3 12b | 12.19 billion | 6.1 | Not available |
| Mistral Nemo 2407 | 12.25 billion | 6.1 | Config supported |
| Phi 3 medium 4k | 13.96 billion | 7.0 | Not available |
| phi 4 | 14.66 billion | 7.3 | Config supported |
| Qwen3 14B | 14.77 billion | 7.4 | Config supported |
| Qwen2.5 14B | 14.77 billion | 7.4 | Config supported |
| Qwen2.5 Coder 14B | 14.77 billion | 7.4 | Config supported |
| DeepSeek R1 Distill Qwen 14B | 14.77 billion | 7.4 | Config supported |
| Mistral Small 24B 2501 | 23.57 billion | 11.8 | Config supported |
| gemma 2 27b | 27.23 billion | 13.6 | Not available |
| gemma 3 27b | 27.43 billion | 13.7 | Not available |
| Qwen3 30B A3B | 30.53 billion | 15.3 | Config supported |
| Qwen3 32B | 32.76 billion | 16.4 | Config supported |
| Qwen2.5 32B | 32.76 billion | 16.4 | Config supported |
| Qwen2.5 Coder 32B | 32.76 billion | 16.4 | Config supported |
| DeepSeek R1 Distill Qwen 32B | 32.76 billion | 16.4 | Config supported |
| Mixtral 8x7B v0.1 | 46.70 billion | 23.4 | Config supported |
| Llama 3.1 70B | 70.55 billion | 35.3 | Not available |
| Llama 3.3 70B | 70.55 billion | 35.3 | Not available |
| DeepSeek R1 Distill Llama 70B | 70.55 billion | 35.3 | Config supported |
| Qwen2.5 72B | 72.71 billion | 36.4 | Config supported |
| Mixtral 8x22B v0.1 | 140.63 billion | 70.3 | Config supported |
| Qwen3 235B A22B | 235.09 billion | 117.5 | Config supported |
| Llama 3.1 405B | 405.85 billion | 202.9 | Not available |
Estimate training cost · Compare LLM API pricing · GPU rental prices