Claude Sonnet 5.5 costs the same per token but up to 30% less per task. Time to re-test your workflow steps
Anthropic shipped Claude Sonnet 5.5, and the interesting number is cost per task, not price per token.
What happened
Anthropic released Claude Sonnet 5.5, the second model in its Claude 5.5 family. Per The Decoder, pricing is unchanged from Sonnet 5 at $2 per million input tokens, $10 per million output tokens and $0.20 for cache reads. Anthropic says the model uses fewer tokens per task, so effective cost drops by up to 30 percent, and output is more than 30 percent faster. It also batches tool calls more often, which cuts the number of steps an agent needs.
On GDPval-AA, a knowledge work benchmark covering 44 professions, Anthropic reports Sonnet 5.5 at 1,844 versus 1,846 for Opus 5.5. On Terminal-Bench 4.0 it jumps from 10.3 to 70.6 percent. One wrinkle: at the "Max" effort setting it scored worse on FrontierCode than at "Xhigh", because it more often split work across sub-agents, causing timeouts or out of scope changes. Simon Willison also hit a Max effort run that burned 128,000 thinking tokens and failed to finish. TechCrunch notes Sonnet 5.5 is the first Sonnet under the same cyber safeguards as Fable and Opus, and Haiku 5.5 is promised in the coming weeks. These are vendor numbers; independent testing still has to confirm them.
My take
If you run Opus on steps that are really well defined work, like summarising a call, drafting a follow up or filling a spreadsheet, this is the week to test Sonnet 5.5 on them. Same token price with fewer tokens per task is a direct cut to your monthly bill, and fewer tool calls means fewer places for a workflow to fail.
How I would do it:
- Pull 20 to 30 real inputs from last month's runs for one step.
- Run them through your current model and Sonnet 5.5 at medium and high effort.
- Compare output quality and total tokens, not the per token price.
- Keep Max effort off by default. More thinking is not always better, and it can blow your budget on a single run.
Keep Opus for the steps that need judgment. Move the routine ones down a tier and let the saved budget pay for better logging.
Matching each step to the right model is one piece of every automation I build for clients, across lead routing, reporting, data entry, content pipelines and agent workflows. Examples live at romielwillautomate.dev.
More posts
- Nvidia's new agent safety platform promises a millisecond quarantine. Could your automations stop a bad run that fast?Sep 29, 2026
- Stop paying an LLM to write a paragraph when you only need a label. Decision models like Jev change the mathSep 29, 2026
- Shopify now lets browser AI agents complete checkout. Your store's data is about to be read by machines firstSep 29, 2026
