Claude effort levels, explained
low, medium, high, xhigh, max —
five levels, one bill, and a decision most teams get wrong in the expensive direction.
Short version: effort caps how many thinking tokens the model may spend before answering. Thinking tokens bill at the output rate. Raise effort only for tasks that actually failed at a lower level.
The five levels
| Level | Behaviour | Typical use | Cost vs low |
|---|---|---|---|
low | Single pass, no visible reasoning | Extraction, classification, formatting, translation | 1× |
medium | Light reasoning | Chat, summarisation, routine analysis | ~3× |
high | Multi-step reasoning | Application code, refactors, debugging | ~8× |
xhigh | Deep, slower reasoning | Hard bugs, architecture, research | ~20× |
max | Largest accepted budget | Benchmark-grade problems | ~45× |
Cost multiples are planning assumptions, not vendor-published numbers. Edit them against your own measured traffic.
Why the output line is the one that moves
A request has two bills: input tokens (your prompt) and output tokens (everything the model
generates). Thinking is generation, so it lands on the output line — and the output rate is
roughly five times the input rate on every current Claude tier. Raising effort from
low to max therefore multiplies the expensive half of your bill
while leaving the cheap half untouched.
Prompt caching is the counterweight: a long, stable system prompt can be read at a fraction of the standard input rate. Cache the stable part, and keep effort low.
Choosing a level, by task shape
- One correct shape of answer. Labels, dates, currency amounts, JSON with a
fixed schema →
low. Reasoning cannot improve a lookup. - Human is waiting. Support chat, autocomplete, anything interactive →
medium. Latency is part of the product. - Output is code or a plan. →
high. This is where most production traffic belongs. - Failure is expensive. Production incidents, migrations, one-off hard
problems →
xhigh. - You are publishing a result. →
max, and usually on a handful of requests, not a queue.
One gotcha: not every level works on every model
Effort levels are per-model capabilities. A model may reject a combination it does not
support — for example, a request that disables thinking entirely while asking for
xhigh or max can come back as an error rather than being silently
downgraded. Validate the level against the model before you ship, and have a fallback so one
bad request does not take down a batch job.
Unsupported effort levels typically fail loudly at request time. Put the level behind a config value, not a hardcoded string, so you can change it without a deploy.
Work out your own number
The calculator on the home page takes your input tokens, output tokens and daily call volume, and shows the per-call and monthly cost at each level and each model. It runs entirely in your browser — nothing is uploaded.