Which effort level should you use — and what will it cost?
Effort is the single biggest lever on a Claude bill that most teams never touch. Pick your task below and get the recommended model and effort level, the cost per call, and the monthly number across Claude Sonnet 5.5, Opus 5.5 and GPT-6 Astra.
1. What are you doing?
This sets a starting recommendation. Everything stays editable.
2. Tune the workload
Defaults are typical for an API-backed app.
Cost
Thinking tokens bill at the output rate, so they land on the output line.
Same workload, every model
Cheapest first. Sorted on your inputs, not on a benchmark score.
| Model | Price / 1M | Per call | Per 30 days |
|---|
Monthly cost at a glance
Same numbers, drawn to scale.
What "effort" actually changes
Effort is a request parameter, not a model setting you buy. It caps how much
thinking the model is allowed to do before it starts writing the answer.
At low you get a single pass — fast, cheap, and fine for anything with one
correct shape of answer. At max the model can chew on the problem for a long
time before committing to a response.
The part that surprises people: thinking tokens are billed as output tokens.
So the effort level does not touch your input cost at all — it moves the output line, and the
output rate is typically 5× the input rate. That is why an unconsidered
max on high-volume traffic is the fastest way to turn a $200/month bill into a
$2,000/month bill.
The rule of thumb: choose the lowest effort level that still produces a correct answer, then raise it only for the specific task that failed. Effort is a per-request decision, not an account-wide one.
Where each level belongs
| Level | Thinking | Reach for it when |
|---|---|---|
low | ×1 | Classification, field extraction, formatting, translation, short replies |
medium | ×3 | Chat bots, summarisation, routine analysis where latency matters |
high | ×8 | Application code, refactors, debugging — the production default |
xhigh | ×20 | Hard bugs, architecture, multi-file migrations, research |
max | ×45 | Benchmark-grade problems. Rarely justified in production |
Multipliers are planning assumptions you can edit with the slider above, not published vendor figures. Measure your own traffic, then set them to match.
Two mistakes that cost real money
-
Running everything at max. Reasoning quality plateaus. Past a point you
are buying latency and tokens, not correctness. Sample 50 real requests at
highand 50 atxhighand compare — most teams find the difference invisible outside genuinely hard tasks. -
Paying for reasoning on deterministic work. If the output is a label, a
date or a JSON object with a fixed schema, effort cannot help you. That is a
lowjob forever.
Frequently asked
Is a higher effort level always more accurate?
No. Accuracy improves steeply at first and then flattens, while cost keeps climbing linearly. For short, well-specified tasks the curve is essentially flat — you pay more and get the same answer.
Which model should I pair with which effort?
Prefer a stronger model at a lower effort over a weaker model at a higher effort for
hard reasoning. Sonnet 5.5 at high handles most production work; reach for
Opus 5.5 when the cost of a wrong answer exceeds the token bill.
Do these prices include prompt caching?
Cached input reads are billed far below the standard input rate, which is why long system prompts are worth caching. The calculator above uses standard input pricing, so treat it as a conservative upper bound.
How current are the numbers?
The data file behind this page records a verification date, shown in the footer. Pricing in this market moves monthly — always confirm against the vendor's own pricing page before you commit a budget.