
Claude Sonnet 5.5 launched on September 28, 2026 at $2 per million input tokens and $10 per million output tokens, the same rates as Sonnet 5 and half of Opus 5.5's $4 and $20. The rate card is only part of the bill. What a task costs depends on how many tokens the model writes, and that depends heavily on the effort setting. Artificial Analysis measured an average of $7.60 per benchmark task at max effort, so before you move real work to a new model, run it against a job you have already done.
The drill problem: a lower price is not a lower bill
You do not want a drill, you want a hole in the wall. A drill that costs half as much but decides to put 50 holes in your plumbing is not a bargain. Models work the same way. A cheaper token only saves money if the model uses a similar number of tokens to finish the job. If it writes a novel when you asked for a paragraph, the total goes up.
Token price versus task cost
A token is a small chunk of text, roughly a word fragment, that the model reads or writes. Input tokens are what you send. Output tokens are what it writes back, and they cost five times as much here ($10 versus $2 per million). That is why wordiness matters: extra output is the expensive part.
Anthropic's launch pitch, as covered by outlets that reported it, is that Sonnet 5.5 costs less per completed task because it needs fewer tokens and steps. Slack's principal engineer Curtis Allen reported roughly 14% fewer output tokens and fewer steps on offline Slackbot evaluations compared with Sonnet 5, with no prompt changes. That is a Slack internal test on constrained work, not a promise for everyone.
Effort levels decide whether you save money
Sonnet 5.5 has five effort levels: low, medium, high, xhigh, and max. They already existed on Sonnet 5, and Anthropic recalibrated them for 5.5. In Anthropic's migration guide, the API defaults to high and the Claude apps default to medium.
Per-task figures reported from Artificial Analysis testing put the spread at about $0.41 at low effort, $0.59 at medium, $1.08 at high, $2.74 at xhigh, and $7.60 at max. Treat the lower figures as reported rather than independently confirmed. The $7.60 number checks out against Artificial Analysis's own methodology: it is a weighted average cost per benchmark task at max effort, around 193,000 output tokens per task. It is not one user's single task, and it sits above both Sonnet 5 at max and Opus 5.5 at max. Sources disagree on the exact Opus 5.5 max figure, so check the current numbers before you quote them.
In practice, max effort without a tight prompt invites the model to overbuild. It adds code, writes extra text, and burns budget. Medium or low is the safer starting point for routine work.
Where Sonnet 5.5 looks strong, and the claims to treat carefully
Reviewers report impressive code demos, and these are single-source reports I could not independently confirm. One creator (the AI Walkthru channel, a small account) reported building a grad school application tracker in under five minutes that passed 23 of 24 checks. Another reviewer showed an interactive 3D earthquake map pulling live data. Code has a clear finish line: it compiles or it does not, and a map either loads or it crashes.
Open-ended writing is different. In one video test, a reviewer asked for a 120-word About page, and the model invented a client statistic, "clients get about five hours a week back," to satisfy the constraints. I could not find that specific test elsewhere, so treat it as an anecdote. The general risk is better supported: truestandard.ai measured Sonnet 5.5's hallucination rate at about the same as Sonnet 5, with roughly 7.4% of DOIs invented in its test. If your name is on the deliverable, you verify every number.
The three-prompt method
Instead of trusting a launch chart, test the model on your own work in three steps.
Pick a real recurring task you already finished. A past client brief or an old newsletter works, as long as you know what the finished product should look like.
Freeze the setup. Use the exact same source files and the exact same prompt as the original. Do not tweak anything to make it easier for the new model.
Run it and score it against your old baseline. Count edits required, check every fact and number, and note whether it invented anything.
The Taskade cost-per-task formula adds a useful lens: cost per task equals cost per attempt multiplied by one divided by your success rate. A cheap attempt that fails half the time costs twice as much per accepted result. Their guidance suggests about 20 of your real tasks, with pass criteria written first.
Which model for which job
If a task has a clear, unambiguous finish line, such as coding a script or editing a document to specific parameters, Sonnet 5.5 is a reasonable default, kept at medium or low effort. If you are brainstorming from scratch or the goals are fuzzy, the flagship Opus 5.5 is the better fit. A common split from the Claude Code crowd is to plan with Opus and implement with Sonnet.
Anthropic's published guidance, as summarized in the research behind this episode, also suggests three system-prompt guardrails: ask for options or an outline before a full build, tell the model to verify facts and prices that may have changed, and tell it to finish all requested items before stopping.
Your homework
Pick one real job you finished recently. Rerun it on the new model with the same files and the same prompt. Score the result against what you made the first time, and check every number it gives you. Then comment which recurring task you would test first.
FAQ
Is Claude Sonnet 5.5 cheaper than Sonnet 5?
Per token, no. Both cost $2 per million input tokens and $10 per million output tokens. Whether a task costs less depends on how many tokens the model uses, which varies with the task and the effort setting. Slack reported about 14% fewer output tokens on its own evaluations.
How much does Claude Sonnet 5.5 cost per task at max effort?
Artificial Analysis measured an average of $7.60 per benchmark task at max effort, around 193,000 output tokens per task. That is an average across its benchmark, not the cost of any single task of yours.
What effort levels does Sonnet 5.5 have?
Five: low, medium, high, xhigh, and max. The API defaults to high and the Claude apps default to medium, according to Anthropic's migration guide.
How do I test a new AI model on my own work?
Pick a real recurring task you already completed, freeze the same files and prompt, run the new model, and score the output against your old result. Check edits needed, accuracy of every number, and any invented facts.
Companion episode:
Want more practical AI breakdowns for small businesses? Get the free playbooks and automations at joebuildsai.com.

