gemma 2 2b VRAM requirements
gemma 2 2b has 2.61 billion parameters in its published weight files. At 4 bits, the weights alone need approximately 1.3 GB. A full inference estimate is unavailable because the published configuration is restricted or uses an attention layout this calculator does not model.
Updated 2026-09-27. Memory estimates use decimal GB. Quantized weights may add scales, metadata and buffers; a compatible implementation and model licence are still required.
Change the memory assumptions
Weights only: 1.3 GB
Memory at different precisions
| Precision | Weights GB | Inference GB | LoRA scenario GB | Full fine-tune GB |
|---|---|---|---|---|
| 16-bit | 5.2 | Not modelled | 7.7 | 49.1 |
| 8-bit | 2.6 | Not modelled | 5.1 | 49.1 |
| 4-bit | 1.3 | Not modelled | 3.8 | 49.1 |
How the estimate is calculated
Weights = parameters × bits ÷ 8. For supported standard attention, KV cache = 2 × layers × KV heads × head dimension × context tokens × batch size × 2 bytes. Inference adds the chosen percentage of runtime headroom to weights and cache. Divide bytes by 1,000,000,000 for GB.
The LoRA scenario keeps the chosen base weights and allows 18 bytes per trainable adapter parameter, plus the activation allowance. Full fine-tuning uses 18 bytes per parameter for a mixed-precision AdamW scenario, plus that allowance. Its total does not fall when you select 4-bit inference weights. These are budget estimates, not measured memory peaks; training context, optimizer, checkpointing and implementation can change them substantially.
Single-GPU candidates for this inference scenario
The offers update with the memory inputs above. Availability not reported means stock has not been confirmed. Memory capacity alone does not establish runtime compatibility or speed.
No GPU recommendation is made without a supported full memory estimate. Compare GPU memory and rental rates.
Estimate the rental cost of a training run · Compare an API bill with self-hosting