Anthropic packed three models into September. Fable 5.1 took the top of the family. Opus 5.5 arrived on 22 September, and six days later, on 28 September, Sonnet 5.5.
On paper it is a familiar line-up: the strongest, the middle one, the fast one. But the most interesting thing about this trio is how far the gap between them has closed.
Every number below comes from Anthropic's own announcements. Details may shift as independent measurements arrive; the direction doesn't look like it will.
Fable 5.1: the model at the top
Fable 5.1 is Anthropic's most capable publicly available model. An interesting detail: it is the same model as Mythos 5.1, only with different safeguards. Mythos 5.1 is not public; it goes to verified organisations in cybersecurity and the life sciences through trusted access programmes.
The highlights:
- Price: 10 dollars per million input tokens, 50 dollars per million output tokens. Cache reads dropped to 0.25 dollars. According to Anthropic, total cost falls by about 25 percent on typical work and by up to 45 percent on agent-heavy work.
- A jump in agentic work: on Terminal-Bench 4.0, which measures coding as an agent, the score rose from 42.0 to 55.8 percent. On Terminal-Bench-Science, which measures scientific work in the terminal, it nearly doubled: from 24.7 to 52.6 percent.
- Cybersecurity: the safety filters now block legitimate work 60 percent less often by mistake. Fable 5.1 can be used to find vulnerabilities in software, but it does not develop attack code for them.
- Science: the announcement has examples from protein design to an elevation map of Venus. Far from my line of work, but it shows where the model is heading.
Opus 5.5: Fable level, lower price
Opus 5.5 is the first model of the new Claude 5.5 family. Anthropic's claim is clear: at the level of Fable 5.1 on most work, it runs 40 percent cheaper than Opus 5 and produces output more than 30 percent faster.
Tokens cost 4 dollars per million input and 20 dollars per million output. When I wrote about Opus 5, those figures were 5 and 25 dollars. The per-token price fell by 20 percent; according to Anthropic, the total cost of running it is 40 percent lower.
There is also a fast mode: up to 2.5 times the speed for twice the price (8 dollars input, 40 dollars output).
In Anthropic's table, Opus 5.5 beats Fable 5.1 on several tests:
- Terminal-Bench 4.0 (coding as an agent): Opus 5.5 at 66.4 percent, Fable 5.1 at 55.8 percent.
- OSWorld 2.1 (computer use): 81.8 against 80.7 percent.
- Humanity's Last Exam (hard reasoning): 67.7 against 65.6 percent.
- AutomationBench (business automation): 40.0 against 31.4 percent.
It is ambitious on safety too. The model was tested before release by external evaluators, METR among them, and achieved the best result to date on Anthropic's automated behavioural audit. Its tendency to try to get around boundaries is 85 percent lower than Opus 5 and Mythos 5.1.
Sonnet 5.5: the new workhorse for everyday work
Sonnet 5.5 costs the same as Sonnet 5: 2 dollars per million input tokens, 10 dollars per million output. But it runs more than 30 percent faster, and according to Anthropic the total cost of most work drops by up to 30 percent.
What really surprises is how close it comes to Opus 5.5:
- Terminal-Bench 4.0: at 70.6 percent, Sonnet 5.5 even beats Opus 5.5 on this test (66.4 percent).
- OSWorld 2.1: 80.1 percent against Opus 5.5's 81.8.
- GDPval-AA (knowledge-work tasks): 1844 against 1846 points. Practically the same.
- Humanity's Last Exam: 64.5 against 67.7 percent. Opus is still ahead on hard reasoning.
Anthropic says Sonnet 5.5 writes more clearly and is strong on everyday work with a clear scope (bug fixing, documents, slides, spreadsheets). And one fun note: it is the first Sonnet to finish Pokémon Red working only from screenshots.
How to read these numbers
I have three caveats.
The numbers are the maker's. All of them come from Anthropic's own tables. Independent measurements may turn out more cautious.
The test versions differ. The Fable 5.1 announcement and the Opus 5.5 announcement use different versions of some tests. Putting numbers side by side across announcements would be misleading; I have taken every comparison from within a single table.
A test is not the job. Scoring 70 percent on a test does not mean 70 percent accuracy on your project. These numbers show a direction; they don't give a guarantee.
Which one where in a production agency
In my work, AI speeds up what happens before and after the shoot: proposals and scope documents, shooting schedules, five-language site content, my own software infrastructure. It doesn't touch the set, the decision or the eye. This trio doesn't change that picture; it changes which job I give to which model.
- Sonnet 5.5: everyday work with a clear scope. Shooting schedules, call sheets, draft proposals, the skeleton of a presentation. Fast and cheap, and no longer merely "good enough".
- Opus 5.5: long and complex work. Site infrastructure, multi-step content work, agent tasks that run for hours. Reaching Fable level at this price was unthinkable a year ago.
- Fable 5.1: the hardest problems. Rare in my work. The top shelf to reach for when something is genuinely stuck.
My default would be Opus 5.5; I drop down to Sonnet 5.5 where speed matters, and I open Fable 5.1 when it is genuinely needed.
The real news: the price-performance curve has broken
The real story this month isn't the individual models. The top performance of the previous generation is moving down to a lower model at a lower price. Opus 5.5 has reached Fable level, and Sonnet 5.5 has almost reached Opus level. The same thing happened the same month on the GPT-6 Astra and GPT-6.1 Sol side, which I wrote about a day earlier.
For a small, busy agency like mine, it means this: work I didn't do a year ago because of the cost now makes sense. Multilingual content, archive tagging, long technical jobs. A powerful model is no longer a luxury; it's an everyday tool.
But let me say it again: the stronger the model, the more convincing its output looks. Looking convincing is not the same as being right. Checking the figure, the sentence, the code is my job, and that job doesn't get handed off.
The ordinary asks for the strongest model. The rebellious knows which job to give to which one.