API models · Side by side

Claude Sonnet 5.5 vs Gemini 3.1 Pro Preview

For text, images and documents, both models are candidates. Audio and video inputs make this a different comparison: Gemini supports those formats here, while Sonnet does not.

The short answer

Choose around the material you need to send. A transcript is text, but a recording is audio; those are different inputs. Gemini also carries a preview designation, which matters when you depend on a stable model version.

Compare the details

Compare context, supported inputs and usage costs. Highlighted rows show a difference.

Claude Sonnet 5.5 vs Gemini 3.1 Pro Preview specifications
FeatureClaude Sonnet 5.5Gemini 3.1 Pro Preview
Input formatsfile, image, textaudio, file, image, text, video
Output formatstexttext
Context / position limit1,000,0001,048,576
Maximum output tokens128,00065,536
Developer featuresReasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Stop sequences, Structured output, Tool selection, Tool calling, Response lengthReasoning output, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, Sampling controls
Cached input / 1M tokens$0.10$0.20
Cache write / 1M tokens$2.50$0.38
USD input / 1M tokens$2.00$2.00
USD output / 1M tokens$10.00$12.00
Model identifieranthropic/claude-sonnet-5.5google/gemini-3.1-pro-preview

API pricing

Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.

Gemini 3.1 Pro Preview: long-prompt rates
  • Long-prompt threshold: 200,000 prompt tokens. Input: $4.00 / 1M. Output: $18.00 / 1M.

Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.

Compare local models

What those prices mean in practice

Each example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.

Token cost examples
WorkloadClaude Sonnet 5.5Gemini 3.1 Pro Preview
1,000 short requestsPer request: 2,000 input + 500 output tokens$9.00$10.00
100 document summariesPer request: 20,000 input + 1,000 output tokens$5.00$5.20
10 long-document requestsPer request: 300,000 input + 2,000 output tokens$6.20$12.36

Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.

Beyond the price tag

Documents versus recordings

A text document or image can be passed to either model. A workflow that sends an audio recording or a video directly needs Gemini’s corresponding input support. Converting that recording into text first would introduce another service, another cost and potentially missing visual or audio information.

Long-document costs

Both offer large context windows. However, a larger accepted input is not automatically an inexpensive request: prompt length can change the pricing tier. The long-document example uses the same input and output sizes for both models so that the effect of the available billing rules is visible.

A preview in production

Keep the exact Gemini preview identifier in the comparison and deployment configuration. A preview and a later stable model should not be treated as an unnamed replacement. For a service with a long support life, include the migration work and model availability in the decision, alongside its token prices.

Common questions

Are these monthly subscription prices?

No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.

Do the cost examples include everything?

They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.