API models · Side by side

GPT-6 Luna vs Gemini 3.8 Flash

Two options for applications that make many requests. Compare the token budget first, then decide whether audio or video support changes the shortlist.

The short answer

Luna covers text, images and files. Gemini Flash also accepts audio and video. A text-only automation should not pay for a different model solely because it supports formats the workflow never uses.

Compare the details

Compare context, supported inputs and usage costs. Highlighted rows show a difference.

GPT-6 Luna vs Gemini 3.8 Flash specifications
FeatureGPT-6 LunaGemini 3.8 Flash
Input formatsfile, image, textaudio, file, image, text, video
Output formatstexttext
Context / position limit1,050,0001,048,576
Maximum output tokens128,00065,536
Developer featuresReasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Structured output, Tool selection, Tool calling, Response lengthReasoning output, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, Sampling controls
Cached input / 1M tokens$0.01$0.08
Cache write / 1M tokens$0.13$0.04
USD input / 1M tokens$0.10$0.75
USD output / 1M tokens$0.50$3.75
Model identifieropenai/gpt-6-lunagoogle/gemini-3.8-flash

API pricing

Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.

GPT-6 Luna: long-prompt rates
  • Long-prompt threshold: 272,000 prompt tokens. Input: $0.20 / 1M. Output: $0.75 / 1M.

Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.

Compare local models

What those prices mean in practice

Each example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.

Token cost examples
WorkloadGPT-6 LunaGemini 3.8 Flash
1,000 short requestsPer request: 2,000 input + 500 output tokens$0.45$3.38
100 document summariesPer request: 20,000 input + 1,000 output tokens$0.25$1.88
10 long-document requestsPer request: 300,000 input + 2,000 output tokens$0.61$2.33

Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.

Beyond the price tag

High-volume text jobs

Classification, short drafting and extraction can involve many small requests. Estimate both prompt and answer size rather than treating one million tokens as one million messages. The short-request example shows the cost of a specified number of calls, which is a useful starting point for a batch of similar tasks.

Multimodal work

Gemini’s audio and video inputs can remove a separate conversion step when the task needs the original recording. Luna can still receive a transcript or extracted frames, but generating those adds work outside this model comparison. Choose based on the information your task needs to retain.

Limits before launch

Use the context and output limits together. A request close to the context ceiling may leave less room for a useful answer. Keep an application-level response budget and consider caching repeated instructions; the base token prices do not cover every optional tool or media charge.

Common questions

Are these monthly subscription prices?

No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.

Do the cost examples include everything?

They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.