GPT-6 Luna
- Input / 1M tokens
- $0.10
- Output / 1M tokens
- $0.50
API models · Side by side
Two options for applications that make many requests. Compare the token budget first, then decide whether audio or video support changes the shortlist.
Luna covers text, images and files. Gemini Flash also accepts audio and video. A text-only automation should not pay for a different model solely because it supports formats the workflow never uses.
Compare context, supported inputs and usage costs. Highlighted rows show a difference.
| Feature | GPT-6 Luna | Gemini 3.8 Flash |
|---|---|---|
| Input formats | file, image, text | audio, file, image, text, video |
| Output formats | text | text |
| Context / position limit | 1,050,000 | 1,048,576 |
| Maximum output tokens | 128,000 | 65,536 |
| Developer features | Reasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Structured output, Tool selection, Tool calling, Response length | Reasoning output, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, Sampling controls |
| Cached input / 1M tokens | $0.01 | $0.08 |
| Cache write / 1M tokens | $0.13 | $0.04 |
| USD input / 1M tokens | $0.10 | $0.75 |
| USD output / 1M tokens | $0.50 | $3.75 |
| Model identifier | openai/gpt-6-luna | google/gemini-3.8-flash |
Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.
Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.
Compare local modelsEach example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.
| Workload | GPT-6 Luna | Gemini 3.8 Flash |
|---|---|---|
| 1,000 short requestsPer request: 2,000 input + 500 output tokens | $0.45 | $3.38 |
| 100 document summariesPer request: 20,000 input + 1,000 output tokens | $0.25 | $1.88 |
| 10 long-document requestsPer request: 300,000 input + 2,000 output tokens | $0.61 | $2.33 |
Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.
Classification, short drafting and extraction can involve many small requests. Estimate both prompt and answer size rather than treating one million tokens as one million messages. The short-request example shows the cost of a specified number of calls, which is a useful starting point for a batch of similar tasks.
Gemini’s audio and video inputs can remove a separate conversion step when the task needs the original recording. Luna can still receive a transcript or extracted frames, but generating those adds work outside this model comparison. Choose based on the information your task needs to retain.
Use the context and output limits together. A request close to the context ceiling may leave less room for a useful answer. Keep an application-level response budget and consider caching repeated instructions; the base token prices do not cover every optional tool or media charge.
No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.
They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.