Why does an LLM need more memory than its weights?
The weight estimate accounts only for parameters at the chosen precision. Inference also needs cache and runtime memory. Training additionally keeps gradients, optimizer state and activations.
The model calculator separates weights from a supported KV-cache estimate and editable runtime headroom. Where the configuration is unavailable or the attention type is not modelled, it does not supply a complete inference estimate.
Use the job cost and fee table, model memory estimates, training calculator, and price trends for the relevant inputs.
Related questions
Numbers on this page come from today's verified snapshot. Full table on the homepage; method in the methodology.