DeepSeek V4.1 Flash
- Input / 1M tokens
- $0.30
- Output / 1M tokens
- $1.20
API models · Side by side
A budget-focused comparison for text and image workflows. Check input and output charges separately before assuming that either model name describes the cheaper option for your requests.
Both accept text and images. Luna also accepts files in this catalog. Match the material you send and the amount of text you need back before selecting a model for high-volume work.
Compare context, supported inputs and usage costs. Highlighted rows show a difference.
| Feature | DeepSeek V4.1 Flash | GPT-6 Luna |
|---|---|---|
| Input formats | image, text | file, image, text |
| Output formats | text | text |
| Context / position limit | 1,048,576 | 1,050,000 |
| Maximum output tokens | 943,718 | 128,000 |
| Developer features | frequency penalty, Reasoning output, logit bias, logprobs, Output limit, min p, presence penalty, Reasoning controls, Thinking effort, repetition penalty, Response format, Seed, Stop sequences, Structured output, Temperature, Tool selection, Tool calling, top k, top logprobs, Sampling controls | Reasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Seed, Structured output, Tool selection, Tool calling, Response length |
| Cached input / 1M tokens | $0.006000 | $0.01 |
| Cache write / 1M tokens | Unknown | $0.13 |
| USD input / 1M tokens | $0.30 | $0.10 |
| USD output / 1M tokens | $1.20 | $0.50 |
| Model identifier | deepseek/deepseek-v4.1-flash | openai/gpt-6-luna |
Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.
Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.
Compare local modelsEach example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.
| Workload | DeepSeek V4.1 Flash | GPT-6 Luna |
|---|---|---|
| 1,000 short requestsPer request: 2,000 input + 500 output tokens | $1.20 | $0.45 |
| 100 document summariesPer request: 20,000 input + 1,000 output tokens | $0.72 | $0.25 |
| 10 long-document requestsPer request: 300,000 input + 2,000 output tokens | $0.92 | $0.61 |
Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.
Small structured responses can make output control important. Specify the fields your application needs and compare the supported response controls. A token rate does not establish classification accuracy, so the budget should describe the intended request rather than claiming a model will produce a particular quality score.
Luna’s file input can be relevant for a document workflow. DeepSeek’s text and image inputs require you to decide how documents reach the model—for example, as extracted text or images. That preparation is part of your application and is not automatically included in a model’s token price.
Multiply the input and output used by one request by the number of requests you expect. Keep caching, media and tools separate unless their charges are known. The examples below show defined text budgets; they are not a prediction of how many tokens your own prompts will require.
No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.
They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.