Sonnet 5.5 edged past Opus. How do you choose a model now?
Claude Sonnet 5.5 scored above Opus 5.5 on one benchmark at half the token price. At some settings it still costs more per task. The published numbers, the independent ones, and an order for choosing models for your own work.

On this page
Anthropic released Claude Sonnet 5.5 on September 28. The number quoted most from the launch was its Terminal-Bench 4.0 score. That test checks whether an agent can finish professional tasks by typing commands into a terminal on its own, and Sonnet 5.5 scored 70.6%, above the larger Opus 5.5 at 66.4%. Its token price is half of Opus 5.5’s.
So is there any reason left to use Opus? Look a little closer at the numbers and the answer gets less simple.
Below, Anthropic’s figures sit next to independent measurements, followed by an order for choosing the right model for your team’s work.
What Anthropic published
Here are the two models side by side on the benchmarks Anthropic reported.
| Benchmark | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% |
| OSWorld 2.1 (partial credit) | 80.1% | 81.8% |
| CursorBench 4.0 | 55.5% | 57.8% |
| GDPval-AA v2.1 | 1844 | 1846 |
| Humanity’s Last Exam (with tools) | 64.5% | 67.7% |
| FrontierCode 1.1 | 46.2% | 54.4% |
Terminal-Bench is the only one Sonnet 5.5 leads. Opus 5.5 stays slightly ahead on the rest, by less than four points everywhere except FrontierCode, where the gap is 8.2.
The jump from the previous generation is large. Sonnet 5 scored 10.3% on the same Terminal-Bench 4.0. The price is unchanged from Sonnet 5 at $2 per million input tokens and $10 per million output, and Anthropic says output is more than 30% faster. It also says the model uses fewer tokens, cutting cost per task by up to 30% against Sonnet 5.
Those are the vendor’s own measurements. When the independent evaluator Artificial Analysis ran Terminal-Bench 4.0 itself, the order held: 64% for Sonnet 5.5 and a little over 60% for Opus 5.5.
A week of new price lists
In the week before Sonnet 5.5, three companies released new models over two days, September 21 and 22. Prices below are base rates per million input and output tokens.
| Model | Released | Input | Output | Compared with the previous model |
|---|---|---|---|---|
| Grok 4.7 | Sep 21 | $2 | $6 | Same as Grok 4.6 |
| Claude Opus 5.5 | Sep 22 | $4 | $20 | 20% below Opus 5 ($5/$25) |
| GPT-6 Sol | Sep 22 | $2 | $10 | Half of GPT-5.6 Sol ($4/$20) |
| GPT-6 Luna | Sep 22 | $0.10 | $0.50 | Half or less of GPT-5.6 Luna ($0.20/$1.20) |
| Claude Sonnet 5.5 | Sep 28 | $2 | $10 | Same as Sonnet 5 |
Sonnet 5.5, GPT-6 Sol and Grok 4.7 now all charge $2 per million input tokens. Input price alone no longer separates them. Some of these models also charge a higher rate for very long prompts, so check the provider’s price list if you work with long documents.
What one task costs
Half the token price sounds like half the cost. What you actually pay depends on how many tokens a model spends to finish a task.
Artificial Analysis runs the tasks behind its own Intelligence Index on each model and publishes the score and the cost per task. By reasoning effort:
| Model and effort | Index score | Cost per task |
|---|---|---|
| Opus 5.5, max | 58 | $5.98 |
| Sonnet 5.5, max | 56 | $7.60 |
| Sonnet 5.5, high | 47 | $1.08 |
| Sonnet 5.5, low | 36 | $0.41 |
At maximum effort, Sonnet 5.5 wrote about 193,000 output tokens per task. That is the most Artificial Analysis has measured for any model, and roughly 60% more than Opus 5.5 at max (about 119,000). So the model with half the token price cost 27% more per task than Opus 5.5, and scored two points lower.
The spread inside one model is wider still. Moving Sonnet 5.5 from low to max multiplies cost per task by more than 18 and lifts the score from 36 to 56. Between high and max alone, cost goes up sevenfold for nine points.
Anthropic’s “up to 30% less per task” depends on the setting too. In Artificial Analysis’s measurements, Sonnet 5.5 at max cost about 50% more per task than Sonnet 5.
The Index tasks are not your tasks, of course. Read the table for direction and make the decision with numbers you measured yourself.
An order for choosing
- Take tasks from last week’s real work. Ten to twenty recurring jobs are enough, such as fixing a failing test or summarizing a customer thread. Write one line per task saying what counts as done.
- Count cost per successful task. Record the tokens each task used and whether it succeeded, then divide total spend by the number of successes. A cheap model that fails twice and succeeds on the third try can cost more than an expensive one that gets it right the first time.
- Tune reasoning effort before switching models. As the measurements above show, one model’s cost can vary more than tenfold with its setting. Run a mid-tier model at medium or high first, and move only the tasks it cannot handle to max or to a larger model.
- Give each kind of work its own default. On the published numbers, Sonnet 5.5 is strong for agents that work by typing commands in a terminal, while Opus 5.5 still leads on hard coding benchmarks like FrontierCode. For high-volume, simple work such as classification or extraction, a small model like GPT-6 Luna is worth testing first. Its input price is one twentieth of Sonnet 5.5’s.
- Re-run the same tasks when a new model ships. Price lists can change several times in one week, as they just did. Keep the task set from step 1 and evaluating a new model takes an hour or two. Pin the model on scheduled work so its cost does not shift without notice.
And in endue
endue gathers models from several providers into one catalog and lets you set the model in three places: the agent’s default in Agent Builder, a single message from the composer, and a routine that pins its own model. The model picker shows the price per million tokens and whether the model supports tool calling, and some models also let you choose a reasoning effort.
That makes the order above easy to carry over. Set a fast, inexpensive model as the agent’s default and reach for a stronger one on the few messages that need it. Pin a model on each routine and scheduled work keeps its cost when you retune the agent.
The settings are covered in Choosing a model, and the models you can pick today, with their prices, are on the models page.
Sources
- Anthropic, Introducing Claude Sonnet 5.5 (September 28, 2026)
- Anthropic, Introducing Claude Opus 5.5 (September 22, 2026)
- Artificial Analysis, Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index, with per-effort results for high and low
- Artificial Analysis, Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index
- OpenAI, API pricing
- xAI, API pricing

