API models · Side by side

DeepSeek V4.1 Flash vs GPT-6 Luna

A budget-focused comparison for text and image workflows. Check input and output charges separately before assuming that either model name describes the cheaper option for your requests.

The short answer

Both accept text and images. Luna also accepts files in this catalog. Match the material you send and the amount of text you need back before selecting a model for high-volume work.

Compare the details

Compare context, supported inputs and usage costs. Highlighted rows show a difference.

DeepSeek V4.1 Flash vs GPT-6 Luna specifications
FeatureDeepSeek V4.1 FlashGPT-6 Luna
Input formatsimage, textfile, image, text
Output formatstexttext
Context / position limit1,048,5761,050,000
Maximum output tokens943,718128,000
Developer featuresfrequency penalty, Reasoning output, logit bias, logprobs, Output limit, min p, presence penalty, Reasoning controls, Thinking effort, repetition penalty, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, top k, top logprobs, Sampling controlsReasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Structured output, Tool selection, Tool calling, Response length
Cached input / 1M tokens$0.006000$0.01
Cache write / 1M tokensUnknown$0.13
USD input / 1M tokens$0.30$0.10
USD output / 1M tokens$1.20$0.50
Model identifierdeepseek/deepseek-v4.1-flashopenai/gpt-6-luna

API pricing

Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.

GPT-6 Luna: long-prompt rates
  • Long-prompt threshold: 272,000 prompt tokens. Input: $0.20 / 1M. Output: $0.75 / 1M.

Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.

Compare local models

What those prices mean in practice

Each example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.

Token cost examples
WorkloadDeepSeek V4.1 FlashGPT-6 Luna
1,000 short requestsPer request: 2,000 input + 500 output tokens$1.20$0.45
100 document summariesPer request: 20,000 input + 1,000 output tokens$0.72$0.25
10 long-document requestsPer request: 300,000 input + 2,000 output tokens$0.92$0.61

Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.

Beyond the price tag

Extraction and classification

Small structured responses can make output control important. Specify the fields your application needs and compare the supported response controls. A token rate does not establish classification accuracy, so the budget should describe the intended request rather than claiming a model will produce a particular quality score.

Files versus extracted text

Luna’s file input can be relevant for a document workflow. DeepSeek’s text and image inputs require you to decide how documents reach the model—for example, as extracted text or images. That preparation is part of your application and is not automatically included in a model’s token price.

Scale the actual request

Multiply the input and output used by one request by the number of requests you expect. Keep caching, media and tools separate unless their charges are known. The examples below show defined text budgets; they are not a prediction of how many tokens your own prompts will require.

Common questions

Are these monthly subscription prices?

No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.

Do the cost examples include everything?

They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.