How does context length affect GPU memory estimates?
In the standard KV-cache calculation used here, doubling context length doubles cache memory while leaving the weight estimate unchanged. Batch size scales that cache estimate too.
The formula uses the publisher’s layer count, key/value heads and head dimension. Sliding-window, compressed-cache and other unsupported attention layouts do not receive a standard-cache estimate.
Use the job cost and fee table, model memory estimates, training calculator, and price trends for the relevant inputs.
Related questions
Numbers on this page come from today's verified snapshot. Full table on the homepage; method in the methodology.