Claude Sonnet 5.5
- Input / 1M tokens
- $2.00
- Output / 1M tokens
- $10.00
API models · Side by side
For text, images and documents, both models are candidates. Audio and video inputs make this a different comparison: Gemini supports those formats here, while Sonnet does not.
Choose around the material you need to send. A transcript is text, but a recording is audio; those are different inputs. Gemini also carries a preview designation, which matters when you depend on a stable model version.
Compare context, supported inputs and usage costs. Highlighted rows show a difference.
| Feature | Claude Sonnet 5.5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Input formats | file, image, text | audio, file, image, text, video |
| Output formats | text | text |
| Context / position limit | 1,000,000 | 1,048,576 |
| Maximum output tokens | 128,000 | 65,536 |
| Developer features | Reasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Stop sequences, Structured output, Tool selection, Tool calling, Response length | Reasoning output, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, Sampling controls |
| Cached input / 1M tokens | $0.10 | $0.20 |
| Cache write / 1M tokens | $2.50 | $0.38 |
| USD input / 1M tokens | $2.00 | $2.00 |
| USD output / 1M tokens | $10.00 | $12.00 |
| Model identifier | anthropic/claude-sonnet-5.5 | google/gemini-3.1-pro-preview |
Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.
Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.
Compare local modelsEach example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.
| Workload | Claude Sonnet 5.5 | Gemini 3.1 Pro Preview |
|---|---|---|
| 1,000 short requestsPer request: 2,000 input + 500 output tokens | $9.00 | $10.00 |
| 100 document summariesPer request: 20,000 input + 1,000 output tokens | $5.00 | $5.20 |
| 10 long-document requestsPer request: 300,000 input + 2,000 output tokens | $6.20 | $12.36 |
Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.
A text document or image can be passed to either model. A workflow that sends an audio recording or a video directly needs Gemini’s corresponding input support. Converting that recording into text first would introduce another service, another cost and potentially missing visual or audio information.
Both offer large context windows. However, a larger accepted input is not automatically an inexpensive request: prompt length can change the pricing tier. The long-document example uses the same input and output sizes for both models so that the effect of the available billing rules is visible.
Keep the exact Gemini preview identifier in the comparison and deployment configuration. A preview and a later stable model should not be treated as an unnamed replacement. For a service with a long support life, include the migration work and model availability in the decision, alongside its token prices.
No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.
They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.