API models · Side by side

Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

Compare two Gemini options for text, documents, audio and video. The supported input formats overlap, while pricing and the preview designation give you different things to consider.

The short answer

Both support the same broad input categories in this comparison. Choose around the task and bill rather than the Flash or Pro name alone. Keep the Pro preview version explicit in a long-lived application.

Compare the details

Compare context, supported inputs and usage costs. Highlighted rows show a difference.

Gemini 3.8 Flash vs Gemini 3.1 Pro Preview specifications
FeatureGemini 3.8 FlashGemini 3.1 Pro Preview
Input formatsaudio, file, image, text, videoaudio, file, image, text, video
Output formatstexttext
Context / position limit1,048,5761,048,576
Maximum output tokens65,53665,536
Developer featuresReasoning output, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, Sampling controlsReasoning output, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, Sampling controls
Cached input / 1M tokens$0.08$0.20
Cache write / 1M tokens$0.04$0.38
USD input / 1M tokens$0.75$2.00
USD output / 1M tokens$3.75$12.00
Model identifiergoogle/gemini-3.8-flashgoogle/gemini-3.1-pro-preview

API pricing

Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.

Gemini 3.1 Pro Preview: long-prompt rates
  • Long-prompt threshold: 200,000 prompt tokens. Input: $4.00 / 1M. Output: $18.00 / 1M.

Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.

Compare local models

What those prices mean in practice

Each example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.

Token cost examples
WorkloadGemini 3.8 FlashGemini 3.1 Pro Preview
1,000 short requestsPer request: 2,000 input + 500 output tokens$3.38$10.00
100 document summariesPer request: 20,000 input + 1,000 output tokens$1.88$5.20
10 long-document requestsPer request: 300,000 input + 2,000 output tokens$2.33$12.36

Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.

Beyond the price tag

Shared input formats

A workflow that accepts video, audio, images and documents can consider either model. That does not establish the same result on every recording or document. Keep the input and requested output identical when reviewing whether another model adds value to a workflow that already functions.

Many short requests or fewer long ones

A small answer and a long report use different output budgets. Compare both sides of the pricing table and include the long-prompt conditions. For scheduled document work, the same token allowance per request provides a clearer budget comparison than a generic monthly spend estimate.

Release stability

The Pro model here is specifically Gemini 3.1 Pro Preview. A preview identifier is not a promise that the endpoint will remain unchanged indefinitely. Track the chosen model in configuration and keep a migration plan if a later release or availability change affects an application you intend to maintain.

Common questions

Are these monthly subscription prices?

No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.

Do the cost examples include everything?

They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.