Claude Sonnet 5.5
- Input / 1M tokens
- $2.00
- Output / 1M tokens
- $10.00
API models · Side by side
Two Claude models with the same supported input types and context allowance, but different usage costs. The decision is whether your particular work benefits enough from Opus to justify the extra spend.
Use the same request size to compare the bill first. Both accept text, images and files, so basic input compatibility does not separate them. Output length and thinking settings can still change what an actual task costs.
Compare context, supported inputs and usage costs. Highlighted rows show a difference.
| Feature | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Input formats | file, image, text | file, image, text |
| Output formats | text | text |
| Context / position limit | 1,000,000 | 1,000,000 |
| Maximum output tokens | 128,000 | 128,000 |
| Developer features | Reasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Stop sequences, Structured output, Tool selection, Tool calling, Response length | Reasoning output, Completion limit, Output limit, Reasoning controls, Thinking effort, Response format, Stop sequences, Structured output, Tool selection, Tool calling, Response length |
| Cached input / 1M tokens | $0.10 | $0.20 |
| Cache write / 1M tokens | $2.50 | $5.00 |
| USD input / 1M tokens | $2.00 | $4.00 |
| USD output / 1M tokens | $10.00 | $20.00 |
| Model identifier | anthropic/claude-sonnet-5.5 | anthropic/claude-opus-5.5 |
Prices are in USD per million tokens. Input is what you send; output is what the model generates. Claude and ChatGPT subscriptions are billed separately.
Tools, images and additional reasoning can add charges. The examples below cover input and output tokens only.
Compare local modelsEach example uses the same token budget for both models. These are calculations, not measured workloads. Cache discounts, media, tools and extra reasoning are excluded.
| Workload | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| 1,000 short requestsPer request: 2,000 input + 500 output tokens | $9.00 | $18.00 |
| 100 document summariesPer request: 20,000 input + 1,000 output tokens | $5.00 | $10.00 |
| 10 long-document requestsPer request: 300,000 input + 2,000 output tokens | $6.20 | $12.40 |
Estimate = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Long-prompt rates are applied where the pricing data provides them. Real requests can use different amounts of output.
A team can keep ordinary drafting and clearly scoped changes on one model while reserving another for selected tasks. That is a routing decision your application makes, not a feature automatically included by choosing Claude. Keep the same documents, instructions and output requirements when deciding whether a switch is useful.
A long answer can cost more than the material sent to the model. For reports and generated code, compare the output rate as well as input. Set an output limit that allows the task to finish, and count any billable reasoning rather than budgeting only the text visible in the final answer.
Both are Anthropic models, but keeping the same vendor does not remove the need to review responses. Tool arguments, JSON fields and the style of generated content can affect downstream behavior. Version and model IDs should be explicit in your configuration so a later update does not quietly change a production workflow.
No. The table compares token charges for using these models in an application. Consumer subscriptions have their own prices and usage rules.
They include input and output token charges under the stated assumptions. Media, tools, additional reasoning and cache operations can add different charges.