Downloadable models · Side by side

Qwen3 4B vs Qwen3 8B

Choose between two sizes in the Qwen3 family. The smaller checkpoint needs less space for its weights; the complete memory requirement also depends on precision, context and runtime.

The short answer

  • Both support thinking and non-thinking modes. Use the correct chat template for the mode you select.
  • Both model cards describe 32,768 native context tokens and a longer context option using YaRN configuration.
  • The 8B checkpoint has roughly twice the reported parameter count. This does not establish twice the quality or a particular token speed.

Compare the details

Compare context, supported inputs and usage costs. Highlighted rows show a difference.

Qwen3 4B vs Qwen3 8B specifications
FeatureQwen3-4BQwen3-8B
Input formatstexttext
Output formatstexttext
Context / position limit32,76832,768
Parameter count4,022,468,0968,190,735,360
Release licenseapache-2.0apache-2.0
Model identifierQwen/Qwen3-4BQwen/Qwen3-8B

API pricing

Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.

Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.

Compare local models

Choosing for your project

Start with the formats you need to send, then compare the context allowance and the cost examples. For an existing application, check tool calls and response formats before changing its model. A higher context limit or a lower token price does not, on its own, tell you which answer will be more useful.

Planning to run it yourself?

Weight size is only part of memory use. Context, concurrency, quantization and runtime overhead matter too. A larger context setting needs separate checks.

What can my computer run?