Model hardware calculator

From model to GPU, priced.

Choose a model and estimate its memory needs. Compare published GPU hardware and rental prices using your precision, context and deployment settings.

Selected Model

bert-base-chinese

google-bert/bert-base-chinese
Parameter count
102,882,442
Architecture
bert
02

Choose your workload

Memory and rental estimate

Nominal storage calculation; selecting a precision does not convert or quantize model files.
Some runtimes may not support each cache format for a given model.
A repository position limit is not automatically treated as your serving context.
Used only in the KV-cache calculation when architecture fields are known.
Multi-GPU and CPU-offload results stay conditional on provider-published hardware and framework details.
Only regions present on a current GPU offer are listed.
Prices in different currencies are not converted or ranked against each other.

Method and limits

How memory and offer checks work

Weights: total reported parameter count × the selected nominal bits per parameter ÷ 8. This is a simplified storage estimate. INT8 and INT4 overhead and unquantized components are not inferred when the repository does not report them. For mixture-of-experts models, all reported parameters remain in the estimate.

KV cache: when published configuration fields support it, layers × requested context × concurrency × KV heads × explicit head dimension × 2 for key and value × selected cache bytes. Missing dimensions leave KV cache unknown; the tool does not infer head dimension from architecture names.

Hosting: only current GPU offerings appear. A lower-bound memory estimate is only a first check. Runtime and system overhead can change the result. We show “meets documented minimum” only when the documented hardware and deployment setup establish it. Unknown provider charges stay unknown.

Accelerate's model estimator builds an empty model from configuration metadata and estimates loading memory. Its documentation explicitly separates model-loading memory from inference memory; this tool follows that boundary and does not treat its own subtotal as total runtime requirements.