Qwen3-4B
- Reported parameters
- 4.02B
- Release license
- apache-2.0
Downloadable models · Side by side
Choose between two sizes in the Qwen3 family. The smaller checkpoint needs less space for its weights; the complete memory requirement also depends on precision, context and runtime.
Compare context, supported inputs and usage costs. Highlighted rows show a difference.
| Feature | Qwen3-4B | Qwen3-8B |
|---|---|---|
| Input formats | text | text |
| Output formats | text | text |
| Context / position limit | 32,768 | 32,768 |
| Parameter count | 4,022,468,096 | 8,190,735,360 |
| Release license | apache-2.0 | apache-2.0 |
| Model identifier | Qwen/Qwen3-4B | Qwen/Qwen3-8B |
Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.
Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.
Compare local modelsStart with the formats you need to send, then compare the context allowance and the cost examples. For an existing application, check tool calls and response formats before changing its model. A higher context limit or a lower token price does not, on its own, tell you which answer will be more useful.
Weight size is only part of memory use. Context, concurrency, quantization and runtime overhead matter too. A larger context setting needs separate checks.
What can my computer run?