I try not to write about the model horse race. These posts are about patterns, practices, and habits that outlive any individual model release. But the Opus 5.5 release provided a good opportunity to demonstrate how you track and evaluate your LLM costs.
Last week, Anthropic released Opus 5.5. They tout a 40% reduction in cost relative to Opus 5. That sounds like a very good deal, but when they report a number like that they’re computing an average over a set of representative workloads. Should you expect to see a 40% reduction in your inference costs?
If you are using evals as a compass, you already have what you need to answer that question. I previously shared an analysis of verbosity in Opus 5 vs. 4.8, and I used that data and code to examine the changes with 5.5.
Data and methods
I have a small collection of twenty-five data science prompts that I use to evaluate the stylistic characteristics of LLM responses. These are designed to be relatively straightforward across a range of task types including factual questions, coding, recommendations, explanation, and text-condensing. I ran each prompt five times against each model: Opus 4.8, Opus 5, and Opus 5.5.1 Reasoning tokens are billed as output tokens, but I track them separately because they aren’t visible as part of the response.
Right on the money
On my workload, Anthropic’s claimed reduction holds up. A median response that costs 5.1 cents of output on Opus 5 costs 3.1 cents with Opus 5.5. That’s right on target with a 40% reduction. Very little of those savings come from the model writing less in its response. The responses themselves are the same length (within noise) as Opus 5’s. The 40% savings breaks down into a few parts. First, the obvious: Opus 5.5’s output price dropped from $25 to $20 per million tokens. Second, we see a notable shift in reasoning: Opus 5.5 spends roughly half the thinking tokens Opus 5 did on the same prompts.
Cheaper than what?
Opus 5.5 is cheaper than Opus 5. But the cost trajectory over the last three months isn’t all downward. When Anthropic released Opus 5, they carried forward Opus 4.8’s pricing structure. But they weren’t equally costly to run: against my workload, Opus 5 billed almost two and a half times the output tokens compared to 4.8. So even if the pricing is fixed, your COGS is moving.
Takeaways
If you were running Opus 4.8 earlier this summer, last week’s drop in unit cost didn’t close that gap. Your unit costs are still up by over a third since June! Of course, pricing changes and new models are released all the time. How well do you understand what those shifts are going to cost you?
At last week’s Copilot launch, which I covered for Every, Satya Nadella called for ‘evals specific to your company,’ because general benchmarks can’t tell you how a model changes outcomes in your organization. A cost quote like Anthropic’s is the same kind of average. The only way to know how model updates impact your workloads is to measure it yourself. A few bits of practical advice:
When you track tokens, break out reasoning and response text separately, even if they are priced the same. There can be invisible shifts in model behavior that you will miss if you simply track in/out.
Price model swaps on your own workloads. You’re already building evals for benchmarking and quality (right?), use them to forecast cost as well.
The code and data used in this analysis are available on GitHub.
Opus 5.5 defaults to a lower reasoning effort than its predecessors. I ran it at both settings, but in this post I’m using the matched-effort runs. At its default it's modestly cheaper still. You can read more in the project repo.


