Claude Opus 5.5 Cost & Pricing
Quick answer
Standard API · USD / 1M tokens- Input tokens
- $4.00/1M
- Output tokens
- $20.00/1M
- Cached input
- $0.20/1M
Claude Opus 5.5 costs $4.00 per 1M input tokens and $20.00 per 1M output tokens at standard API rates. That is 20% cheaper per input or output token than Opus 5, while cached reads are 60% cheaper at $0.20 per 1M. Fast mode doubles the standard rate to $8.00 / $40.00, and Batch halves it to $2.00 / $10.00.
The biggest variable in your real bill is not the unit price — it is your cache hit rate.Calculate your cost
Full price table
Six ways to price a requestSwipe horizontally to see output and usage →
| Rate tier | Input | Output | Use this when |
|---|---|---|---|
| Standard | $4.00/1M | $20.00/1M | Default API requests |
| Cache read | $0.20/1M | — | Previously written prompt prefix |
| Cache write | $5.00/1M | — | First write of a reusable prefix |
| Batch Derived | $2.00/1M | $10.00/1M | Asynchronous, non-urgent work |
| Fast mode Derived | $8.00/1M | $40.00/1M | Latency-sensitive requests |
| US-only Derived | $4.40/1M | $22.00/1M | US-only inference requirement |
Derived rates apply Anthropic's published Batch 50% discount, Fast 2×, and US-only 1.1× multipliers to the standard rates. Cache write here refers to the listed five-minute tier.
Cost calculator
Calculate one call, then project the same workload over a month.
- Input charge $4.00 /1M
- $0.40
- Output charge $20.00 /1M
- $0.20
- Monthly calls
- 10
“$8 / $40” is not the standard price
Those numbers are for Fast mode, a derived 2× rate that can deliver up to 2.5× faster output. The standard API price remains $4.00 input and $20.00 output per 1M tokens. Anthropic says Fast mode is available in Claude Code and on the Claude Platform; check its current availability for your account before budgeting for it.
Worked examples
The arithmetic, line by lineAssume 1M input tokens and 100K output tokens for the month. No cache writes or platform features are included.
1 × $4.00 + 0.1 × $20.00 = $6.000.1 × $4.00 + 0.9 × $0.20 + 0.1 × $20.00 = $2.581 × $2.00 + 0.1 × $10.00 = $3.001 × $5.00 + 0.1 × $25.00 = $7.500.1 × $5.00 + 0.9 × $0.50 + 0.1 × $25.00 = $3.45Cache writes cost $5.00 per 1M tokens, so the cache saves money only when that prefix is reused enough times.
Opus 5.5 vs Opus 5 vs Sonnet 5.5
Swipe horizontally to compare rates →
| Model | Input | Output | Cache read | Positioning |
|---|---|---|---|---|
| Claude Opus 5.5 | $4.00/1M | $20.00/1M | $0.20/1M | Everyday driver for agentic coding and enterprise work |
| Claude Opus 5 | $5.00/1M | $25.00/1M | $0.50/1M | Previous Opus generation |
| Claude Sonnet 5.5 | $2.00/1M | $10.00/1M | $0.20/1M | High-performance coding and agent model; half the Opus 5.5 standard price |
All prices in this comparison are USD per 1M tokens.
How to lower your bill
- Reuse stable prompt prefixes. Compare the available cache TTL tiers against your reuse interval; count cache writes as well as reads.
- Send non-urgent jobs through Batch, which is priced at half the standard input and output rates.
- Route simpler tasks to Haiku 4.5 at $1.00 input / $5.00 output per 1M tokens, or to Sonnet, and keep Opus for work that needs it.
- Control thinking and effort levels where the model API allows it, then measure both answer quality and output tokens.
When caching costs you money
A five-minute cache write costs $5.00 per 1M tokens; a cached read costs $0.20, versus $4.00 uncached. The first call costs $1 more; each later call saves $3.80. Use the prefix at least twice within the 5-minute TTL (the write plus one read) to break even. If calls are more than about five minutes apart, do not write this five-minute cache.
Turn thinking down on routine calls
Higher effort adds thinking tokens billed at $20.00 per 1M output tokens—the largest avoidable line item on routine calls. Drop effort to low or medium for extraction, formatting and short rewrites.
Frequently asked questions
Seven pricing questionsHow much does Claude Opus cost?
Claude Opus 5.5 costs $4.00 per 1M input tokens and $20.00 per 1M output tokens at standard API rates. A workload with 1M input and 100K output tokens costs $6.00 before cache writes or other features.
Is Opus 5.5 cheaper than Opus 5?
Yes. Opus 5.5 standard input and output rates are 20% lower per token than Opus 5, and cached reads are 60% lower. For 1M input and 100K output tokens without caching, that is $6.00 versus $7.50.
How much more expensive is Opus Fast Mode?
Fast mode is a derived 2× tier: $8.00 input and $40.00 output per 1M tokens, compared with $4.00 and $20.00 at standard rates. Anthropic says it can be up to 2.5× faster.
How much does Batch save?
Batch is a derived 50% discount on standard input and output rates. Opus 5.5 Batch costs $2.00 input and $10.00 output per 1M tokens; the worked example falls from $6.00 to $3.00.
How much do Claude cache tokens cost?
Opus 5.5 cache reads cost $0.20 per 1M tokens and five-minute cache writes cost $5.00 per 1M. With a 90% hit rate on the worked example, the read and standard token charges total $2.58 before cache writes.
Opus 5.5 or Sonnet 5.5?
Sonnet 5.5 costs $2.00 input and $10.00 output per 1M tokens, exactly half the Opus 5.5 standard rates. Use task-level quality and latency measurements to decide which model earns its cost.
Does US-only inference cost more?
Yes. Anthropic lists US-only inference at 1.1× standard input and output rates. For Opus 5.5 that derives to $4.40 input and $22.00 output per 1M tokens.
Compare other models
5 verified models · USD per 1M tokens| Claude Haiku 4.5claude-haiku-4-5 | Anthropic | $1.00/1M | $5.00/1M | $0.10/1M | $1.25/1M | 200K |
| Claude Sonnet 5.5claude-sonnet-5-5 | Anthropic | $2.00/1M | $10.00/1M | $0.20/1M | $2.50/1M | 1M |
| Claude Opus 5.5claude-opus-5-5 | Anthropic | $4.00/1M | $20.00/1M | $0.20/1M | $5.00/1M | 1M |
| Claude Opus 5claude-opus-5 | Anthropic | $5.00/1M | $25.00/1M | $0.50/1M | $6.25/1M | 1M |
| Claude Fable 5.1claude-fable-5-1 | Anthropic | $10.00/1M | $50.00/1M | $0.25/1M | $12.50/1M | 1M |
USD per 1M tokens. Base API rates. Tap a heading to sort; scroll for more columns on narrow screens.