Claude effort levels, explained

low, medium, high, xhigh, max — five levels, one bill, and a decision most teams get wrong in the expensive direction.

Short version: effort caps how many thinking tokens the model may spend before answering. Thinking tokens bill at the output rate. Raise effort only for tasks that actually failed at a lower level.

The five levels

LevelBehaviourTypical useCost vs low
lowSingle pass, no visible reasoningExtraction, classification, formatting, translation1×
mediumLight reasoningChat, summarisation, routine analysis~3×
highMulti-step reasoningApplication code, refactors, debugging~8×
xhighDeep, slower reasoningHard bugs, architecture, research~20×
maxLargest accepted budgetBenchmark-grade problems~45×

Cost multiples are planning assumptions, not vendor-published numbers. Edit them against your own measured traffic.

Why the output line is the one that moves

A request has two bills: input tokens (your prompt) and output tokens (everything the model generates). Thinking is generation, so it lands on the output line — and the output rate is roughly five times the input rate on every current Claude tier. Raising effort from low to max therefore multiplies the expensive half of your bill while leaving the cheap half untouched.

Prompt caching is the counterweight: a long, stable system prompt can be read at a fraction of the standard input rate. Cache the stable part, and keep effort low.

Choosing a level, by task shape

One gotcha: not every level works on every model

Effort levels are per-model capabilities. A model may reject a combination it does not support — for example, a request that disables thinking entirely while asking for xhigh or max can come back as an error rather than being silently downgraded. Validate the level against the model before you ship, and have a fallback so one bad request does not take down a batch job.

Unsupported effort levels typically fail loudly at request time. Put the level behind a config value, not a hardcoded string, so you can change it without a deploy.

Work out your own number

The calculator on the home page takes your input tokens, output tokens and daily call volume, and shows the per-call and monthly cost at each level and each model. It runs entirely in your browser — nothing is uploaded.

Ad slot · 970×250 · in-content

Related