Sept 28, 2026  ·  Anthropic launch

Claude Sonnet 5.5 is out.

Here are the 10 things that matter.

I went through the whole announcement, every chart, and what people built on day one. Here's what's actually new.

What we're covering today
  1. 01A race between the old Sonnet and the new one
  2. 02How much cheaper every job gets
  3. 03How much you have to babysit Sonnet
  4. 04How close Sonnet gets to Opus
  5. 05Sonnet is good at gaming now
  6. 06Sonnet is a design genius
  7. 07What people built with Sonnet on day one
  8. 08Every score Anthropic published
  9. 09Where Opus still wins
  10. 10How to switch to Sonnet today
01  ·  The race

Old Sonnet vs new Sonnet. Same request.

Anthropic gave both models the exact same request. Pick one below and hit start. Both models start at the same moment.

The request
Old Sonnet 5
New Sonnet 5.5

Each side thinks first, then writes the code, then runs it. Watch the number at the top of each side. That's how many tokens it used. Tokens are how AI usage gets measured, so fewer tokens means your Claude plan lasts longer.

02  ·  The price

Same price. But every job costs less.

The price didn't change. The new Sonnet just needs fewer tokens to finish the same job, so each job ends up cheaper. Anthropic says up to 30% cheaper, and over 30% faster.

If a job cost you $10 before
Old Sonnet
$10
→
New Sonnet
$7
If an answer took 10 seconds before
Old Sonnet
10s
→
New Sonnet
7.5s
A real example

A finance company called Balyasny tested it on 2,441 real tasks. The old Sonnet used about 497,000 tokens per answer. The new Sonnet used about 121,000. That's about 4 times less, and the answers were better.

If you're on a Claude plan, this means your usage goes further before you hit your limit.

03  ·  Doing the whole job

Sonnet finishes 7 out of 10 jobs on its own. The old Sonnet finished 1.

This test is called Terminal-Bench. The AI gets a real job, like "build this" or "fix this." Nobody helps it. It only counts if the end result actually works.

Real jobs from the test
Fix a bank's alert system

Customers have to get their "low balance" alert within 5 seconds. If the alert shows up late, the job fails.

Fix a broken sign-up form

Every new lead has to be checked and saved correctly, with no duplicates.

Speed up a slow website

Pages have to load under a time limit, and nothing is allowed to break.

Audit medical insurance claims

Decide which claims to approve or reject, like a real claims auditor would.

A skilled person would need hours for each one. A checker tests the result at the end. It passes only if it works.

Out of 10 real jobs, how many did it finish?
Old Sonnet 5
1 of 10
→
New Sonnet 5.5
7 of 10

It even edged out Opus 5.5, the bigger model that costs twice as much. Opus finished about 6.6 out of 10.

Both scores use each model's highest effort setting

What early testers noticed
It stops to ask you questions less.

An app builder called Base44 said it rarely stopped in the middle of a build to wait for an answer. So fewer builds get stuck.

It does several steps at once.

The old Sonnet did things one at a time. The new one groups them together, so it finishes in fewer steps.

It figures out your project fast.

Testers said it quickly understood big existing projects, and kept going on tasks that ran for hours.

Anthropic's chartBlue = new Sonnet Orange = Opus Pale green = old Sonnet
Terminal-Bench 4.0 score vs cost chart from Anthropic
  • Higher up means it finished more jobs.
  • Further right means each job cost more money.
  • Find "Med" on the blue line. That's the default setting. Each job costs under $1, and it still beats the old Sonnet's best score, which cost over $10 a job.
04  ·  Sonnet vs Opus

Sonnet basically ties Opus. And Sonnet costs half as much.

Opus is Anthropic's bigger model, and it costs twice as much to use. In this test, experts looked at real work from 44 different jobs, like reports, spreadsheets and plans. Then they picked which AI did it better.

Opus 5.5
1846
New Sonnet 5.5
1844
Old Sonnet 5
1449
Opus scored 1846. The new Sonnet scored 1844. That's a 2-point difference.
Same story on a second office testBlue = new Sonnet Orange = Opus Pale green = old Sonnet
AA-Briefcase score vs cost chart from Anthropic
  • The blue and orange lines almost sit on top of each other. For the same money, you get about the same quality.
  • The pale green line is the old Sonnet. Look how far below it is. That gap is the upgrade.
What companies saw when they tested it

Base44, an app builder, tested it on 118 apps. It needed 3.6 tries per app. Opus 5 needed 7.7.

Zendesk ran it on hundreds of real support tickets. Tickets got solved 20% faster.

Box said it double-checks facts against the original documents, and it caught mistakes the old Sonnet missed.

05  ·  It can see

Sonnet beat Pokémon Red just by looking at the screen.

It only saw screenshots of the game, and it still beat it. No Sonnet model has done that before. It also got way better at reading charts.

Sonnet 5.5 in first place in a live Pokémon raceThis streamer races 4 AIs through Pokémon Yellow at the same time. Sonnet 5.5 is top left, and it's leading.
Sonnet 5.5 invented its own Pokémon-style gameHe asked for one, and Sonnet made "Nahual", a Game Boy game with Mexican monsters and real battles.
Show it 100 charts. How many does it read correctly?
Old Sonnet 5
16
→
New Sonnet 5.5
62

Why this matters for you: send it a screenshot of an error, a chart, or a photo of your notes, and it actually understands it.

06  ·  Design

Sonnet made its own launch video.

Someone asked Sonnet 5.5 to make a video about its own release. This is what came back: animated charts, glowing numbers, and sound design.

Sonnet 5.5 made a launch video about itself.28 seconds of animated charts and sound. He says Sonnet built it about 2x faster than Opus 5.5.

Anthropic saw the same thing in their own tests. They gave it a company's earnings report and a slide template, and asked for a 10-slide deck. Two experts said the first draft was ready to send with no edits.

It also writes more clearly than the old Sonnet. Testers said it feels more like working with a partner.

A Roman soldier, built in 3D.He asked Sonnet 5.5 to make it in Blender, a free 3D design app.

“Claude Sonnet 5.5 cooks.”

Tyler Nishida, designer at Every

07  ·  Day one

What people built with Sonnet in the first few hours.

A tiny emoji, turned into a real toyHe gave Sonnet a tiny pixel emoji and asked for a 3D model. A few hours later, it was printed and sitting on his desk.
A zombie game from one promptA game you can play right in your browser. His prompt ended with "don't ask questions, ship it."
The same tree, old vs newSame request to both models. Watch the difference.
A whiteboard app that looks hand-drawnA real working app. You drag shapes, connect them, and doodle on top.
Sunlight in a forestThis isn't a video. The picture and the sound are both made by code, live in the browser.
All of dinosaur history in one minuteAn animated story from one prompt, old Sonnet vs new Sonnet.
08  ·  The full scorecard

Every score Anthropic published.

Blue is the new Sonnet. Orange is Opus. The highlighted score won that row. Opus still wins most rows, but look how close the new Sonnet gets, and how far it jumped from the old one.

New Sonnet 5.5Old Sonnet 5Opus 5.5GPT-6 Sol
Doing a whole coding job aloneTerminal-Bench 4.070.6%★▲ +60 pts10.3%66.4%—
Code a developer would approveFrontierCode 1.152.1%▲ +10 pts42.4%54.4%49.3%
Real coding tasksCursorBench 4.055.5%▲ +21 pts34.1%57.8%—
Real office work from 44 jobsGDPval-AA1844▲ +395144918461487
Long office projectsAA-Briefcase1811▲ +452135918221483
Really hard expert questionsHumanity's Last Exam64.5%▲ +10 pts54.9%67.7%—
Using a computer like a personOSWorld 2.180.1%▲ +23 pts57.0%81.8%—
Reading chartsChartography61.6%▲ +46 pts15.6%64.4%53.6%
Higher is better on every row. ★ = top score in that row. ▲ = how much the new Sonnet went up from the old one. Dashes mean OpenAI didn't publish a score. Source: Anthropic.

And here's what each model costs.

Prices are per million tokens. "In" is what you send Claude. "Out" is what Claude writes back. The new Sonnet costs exactly what the old one did.

ModelInOutCompared to Sonnet 5.5
Claude Fable 5.1The biggest, for the hardest work$10$505× the price
Claude Opus 5.5Long projects and hard problems$4$202× the price
Claude Sonnet 5.5New today. Speed and smarts together$2$10★ baseline
Claude Sonnet 5The old Sonnet$2$10Same price
Claude Haiku 4.5The fastest and cheapest$1$5Half the price
Source: Anthropic's pricing page. Haiku 5.5 is coming in the next few weeks.
09  ·  The catch

Where Opus still wins.

Big, messy projects.

Anthropic says it plainly: when a project needs a lot of careful thinking, Opus 5.5 is still better.

Max effort can be really slow.

One tester joked he'd die of old age before his Max effort build finished.

Risky security requests get sent to the old model.

If you ask for hacking-type stuff, it switches back to the old Sonnet. Normal coding and bug fixing aren't affected.

Use the new Sonnet for
  • Fixing bugs and small changes
  • Docs, slides and spreadsheets
  • Quick back and forth when you want speed
  • Lots of tasks when you're watching your limits
Use Opus for
  • Big projects from scratch
  • Hard problems that need careful thinking
  • Planning how something should be built
One more thing Anthropic tested

Claude runs inside a locked-down space when Anthropic tests it. They check if a model tries to break out. Sonnet 5.5 was the least likely of all their models to even poke at the walls.

“When Opus 5.5 sets the plan for a game, I would feel confident letting Sonnet 5.5 build it.”

Kevin Ngo, creative coder

10  ·  Use it today

How to switch to Sonnet right now.

1

Pick it in the model menu

In the Claude app, choose Sonnet 5.5 from the model menu. In Claude Code, type /model, press Enter, and choose Sonnet 5.5.

2

Leave effort on Medium

Effort is how hard Claude thinks before it answers. Think of it like a dial. Medium is the default, and it's enough for most jobs.

3

Use your free usage reset

If you're on Pro, Max or Team, Anthropic gave you one free usage reset. You can use it any time before Oct 22.

4

Watch for Haiku 5.5

Haiku is the smallest and cheapest model. The 5.5 version is coming in the next few weeks.

Claude Code Club

Want to actually build with Sonnet 5.5?

Inside the Club, we show you step by step how to build websites, apps and automations with Claude Code. No coding background needed.

Learn Claude Code. Earn income. Only $9.

Prefer to read? Here's the full write-up on the blog.

Sources