Does four-bit quantization use a quarter of 16-bit memory?
For the weight arithmetic alone, four bits is one quarter of sixteen bits per parameter. Total application memory does not necessarily fall by the same fraction, because caches, metadata and runtime allocations remain.
Compare weight-only figures separately from the inference estimate. Quantized formats can introduce additional metadata, and a runtime must support the chosen model and format.
Use the job cost and fee table, model memory estimates, training calculator, and price trends for the relevant inputs.
Related questions
Numbers on this page come from today's verified snapshot. Full table on the homepage; method in the methodology.