Here are the 10 things that matter.
I went through the whole announcement, every chart, and what people built on day one. Here's what's actually new.
Anthropic gave both models the exact same request. Pick one below and hit start. Both models start at the same moment.
Each side thinks first, then writes the code, then runs it. Watch the number at the top of each side. That's how many tokens it used. Tokens are how AI usage gets measured, so fewer tokens means your Claude plan lasts longer.
The price didn't change. The new Sonnet just needs fewer tokens to finish the same job, so each job ends up cheaper. Anthropic says up to 30% cheaper, and over 30% faster.
A finance company called Balyasny tested it on 2,441 real tasks. The old Sonnet used about 497,000 tokens per answer. The new Sonnet used about 121,000. That's about 4 times less, and the answers were better.
If you're on a Claude plan, this means your usage goes further before you hit your limit.
This test is called Terminal-Bench. The AI gets a real job, like "build this" or "fix this." Nobody helps it. It only counts if the end result actually works.
Customers have to get their "low balance" alert within 5 seconds. If the alert shows up late, the job fails.
Every new lead has to be checked and saved correctly, with no duplicates.
Pages have to load under a time limit, and nothing is allowed to break.
Decide which claims to approve or reject, like a real claims auditor would.
A skilled person would need hours for each one. A checker tests the result at the end. It passes only if it works.
It even edged out Opus 5.5, the bigger model that costs twice as much. Opus finished about 6.6 out of 10.
Both scores use each model's highest effort setting
An app builder called Base44 said it rarely stopped in the middle of a build to wait for an answer. So fewer builds get stuck.
The old Sonnet did things one at a time. The new one groups them together, so it finishes in fewer steps.
Testers said it quickly understood big existing projects, and kept going on tasks that ran for hours.
Opus is Anthropic's bigger model, and it costs twice as much to use. In this test, experts looked at real work from 44 different jobs, like reports, spreadsheets and plans. Then they picked which AI did it better.
Base44, an app builder, tested it on 118 apps. It needed 3.6 tries per app. Opus 5 needed 7.7.
Zendesk ran it on hundreds of real support tickets. Tickets got solved 20% faster.
Box said it double-checks facts against the original documents, and it caught mistakes the old Sonnet missed.
It only saw screenshots of the game, and it still beat it. No Sonnet model has done that before. It also got way better at reading charts.
Why this matters for you: send it a screenshot of an error, a chart, or a photo of your notes, and it actually understands it.
Someone asked Sonnet 5.5 to make a video about its own release. This is what came back: animated charts, glowing numbers, and sound design.
Anthropic saw the same thing in their own tests. They gave it a company's earnings report and a slide template, and asked for a 10-slide deck. Two experts said the first draft was ready to send with no edits.
It also writes more clearly than the old Sonnet. Testers said it feels more like working with a partner.
“Claude Sonnet 5.5 cooks.”
Tyler Nishida, designer at Every
Blue is the new Sonnet. Orange is Opus. The highlighted score won that row. Opus still wins most rows, but look how close the new Sonnet gets, and how far it jumped from the old one.
| New Sonnet 5.5 | Old Sonnet 5 | Opus 5.5 | GPT-6 Sol | |
|---|---|---|---|---|
| Doing a whole coding job aloneTerminal-Bench 4.0 | 70.6%★▲ +60 pts | 10.3% | 66.4% | — |
| Code a developer would approveFrontierCode 1.1 | 52.1%▲ +10 pts | 42.4% | 54.4% | 49.3% |
| Real coding tasksCursorBench 4.0 | 55.5%▲ +21 pts | 34.1% | 57.8% | — |
| Real office work from 44 jobsGDPval-AA | 1844▲ +395 | 1449 | 1846 | 1487 |
| Long office projectsAA-Briefcase | 1811▲ +452 | 1359 | 1822 | 1483 |
| Really hard expert questionsHumanity's Last Exam | 64.5%▲ +10 pts | 54.9% | 67.7% | — |
| Using a computer like a personOSWorld 2.1 | 80.1%▲ +23 pts | 57.0% | 81.8% | — |
| Reading chartsChartography | 61.6%▲ +46 pts | 15.6% | 64.4% | 53.6% |
Prices are per million tokens. "In" is what you send Claude. "Out" is what Claude writes back. The new Sonnet costs exactly what the old one did.
| Model | In | Out | Compared to Sonnet 5.5 |
|---|---|---|---|
| Claude Fable 5.1The biggest, for the hardest work | $10 | $50 | 5× the price |
| Claude Opus 5.5Long projects and hard problems | $4 | $20 | 2× the price |
| Claude Sonnet 5.5New today. Speed and smarts together | $2 | $10 | ★ baseline |
| Claude Sonnet 5The old Sonnet | $2 | $10 | Same price |
| Claude Haiku 4.5The fastest and cheapest | $1 | $5 | Half the price |
Anthropic says it plainly: when a project needs a lot of careful thinking, Opus 5.5 is still better.
One tester joked he'd die of old age before his Max effort build finished.
If you ask for hacking-type stuff, it switches back to the old Sonnet. Normal coding and bug fixing aren't affected.
Claude runs inside a locked-down space when Anthropic tests it. They check if a model tries to break out. Sonnet 5.5 was the least likely of all their models to even poke at the walls.
“When Opus 5.5 sets the plan for a game, I would feel confident letting Sonnet 5.5 build it.”
Kevin Ngo, creative coder
In the Claude app, choose Sonnet 5.5 from the model menu. In Claude Code, type /model, press Enter, and choose Sonnet 5.5.
Effort is how hard Claude thinks before it answers. Think of it like a dial. Medium is the default, and it's enough for most jobs.
If you're on Pro, Max or Team, Anthropic gave you one free usage reset. You can use it any time before Oct 22.
Haiku is the smallest and cheapest model. The 5.5 version is coming in the next few weeks.
Inside the Club, we show you step by step how to build websites, apps and automations with Claude Code. No coding background needed.
Learn Claude Code. Earn income. Only $9.Prefer to read? Here's the full write-up on the blog.