AI & LLM hosting
How does context length change inference memory needs?
Short answer
Longer context needs more memory on top of the model weights, mainly for the attention state kept for each token, and it grows with concurrent requests. The size depends on the model architecture and the runtime, so a single figure does not fit every model.
What changes the answer
- Model architecture and how the runtime stores attention state
- Context length and number of parallel requests
- Runtime allocation settings
Weights are only one part of the budget
The weights are the fixed cost of loading a model. A simple estimate multiplies the parameter count by the bytes used per parameter, as the Accelerate memory estimator does for model weights. That figure does not include everything the runtime needs while generating text.
What context adds
While generating, a transformer model keeps intermediate attention state for the tokens already in the conversation so it does not recompute them. The more tokens in the prompt and reply, the more of this state must be held in memory. How much memory each token needs depends on the model's architecture, such as the number of layers and attention heads, and on the precision the runtime uses. For that reason two models with the same parameter count can need quite different amounts of memory at the same context length.
Concurrency multiplies it
Each active request has its own context. Ollama's FAQ says required RAM scales with the number of parallel requests multiplied by the context length. A setup that loads and answers a single short prompt may fail when several long prompts arrive together. Decide what context length and concurrency you really need, then test at those values.
What to do with this
- Treat the weight estimate as a lower bound for planning, not as a guarantee that the model will run.
- Check the runtime's documentation for how it sets context length and parallelism.
- Leave headroom for the operating system and other processes sharing the memory.
What HostCritiq can and cannot say
The memory estimator is a planning aid built on stated assumptions. Missing architecture metadata keeps its result conditional. It does not measure speed, and it cannot confirm that a specific model will load on a specific rented server.
A way to reason about it
Suppose you plan one request at a time with a short context, then double the context and allow two parallel requests. Under the scaling Ollama states, the part of the memory that scales with context and parallelism can grow fourfold, while the weights stay the same. Treat that as the shape of the risk, not a measured figure for your model.