Video walkthroughs · Claude Code Club
TL;DR
Run two models on the same six one-shot jobs, hide which side is which, and score the output before you look at the bill. In our run Sonnet 5.5 won five of six tests and finished in 57 minutes against 94 for GPT-6.1 Sol, while Sol cost $2.64 against $6.68. Cost per task is not cost per token, so pick the model by the job and not by the price sheet.
Most model comparisons you read online are one person's vibe after a few chats. A better way exists, and it fits in an afternoon. You give two models the same short brief, you hide which one made which result, you pick the winner with your own eyes, and only then do you look at the clock and the bill. This guide walks through that method and uses our own six-test run of Claude Sonnet 5.5 against GPT-6.1 Sol as the worked example.
Watch the full head-to-head: six use cases, blind A/B, real time and cost numbers.
A head-to-head without a question produces a pile of screenshots and no decision. Before you run anything, write one sentence that starts with 'I am choosing a model for'. For us it was 'building websites, games and other visual things from a short prompt with no scaffolding'. That sentence decides which tests are worth running and which are a waste of an hour.
There are two honest kinds of question. The first is 'which model should be my default', which needs a spread of jobs. The second is 'which model should run this one recurring task', which needs one test repeated a few times. This guide covers the first kind, because it is the harder one to get right, and the second kind is just a smaller version of the same process.
Write the question down where you can see it while you score. The question is the judge, and your mood on the day is not. When a result is impressive but has nothing to do with your question, it does not count.
Easy tasks make every model look the same. If both models can do the job in their sleep, you learn nothing. The tests worth running are the ones where the ceiling is high and the results can be compared by eye in under a minute. We used six:
Look at what those six have in common. Every one of them ends in something you can open and judge immediately. Every one has a wide range of possible quality, so a gap between models shows up instead of hiding. And every one mixes several skills at once, such as design, code, interaction and writing, so a model has to be good at the whole job and not at one trick.
Each brief was short and open. Each one told the model what to make and what a good result feels like, then left every choice to the model. Brand, genre, subject, palette and structure were all up to the model, which is exactly why the results separate. The exact wording of our six prompts lives inside the Claude Code Club, and the decision to keep them short is the part you can copy today.
We gave both models very little. No skills, no reference images, no design system, no example code, and no follow-up turns. One prompt each, one attempt each. That is what we mean by one-shot, and it is the most important rule in the whole method.
The reason is simple. The moment you add skills, references or coaching, you are measuring how well each model uses your help, and the help itself becomes a hidden variable. If you want to know which model has better taste and judgment by default, you have to take the scaffolding away. You can test a full setup later as a second round.
The other half of the rule is symmetry. Both models get the identical text, the same tools and the same empty starting folder. If one model needs a different harness or an extra instruction to even start, write that down as a finding, because it counts against ease of use. Do not quietly fix it for one side.
People are bad at judging work once they know who made it. If you believe one model is better, you will forgive its rough edges and notice the other one's. Blinding fixes this cheaply. Have each model write its output into a folder labeled only A or B, let a coin flip or a script decide which model goes in which folder, and keep the key somewhere you will not look until you have scored all six tests.
In our run we did not know which side was which until we had picked a winner for each test. On the landing page test, for example, the left side was clearly richer, with parallax, a mouse-following glow, and a page that moved from morning to night as you scrolled. We picked it before seeing whose it was. Only then did the reveal show it was Sonnet 5.5, and the other side was Sol.
If you are testing alone, the simplest blind setup is to ask a second Claude Code session to set it up for you. Tell it to run both models, save each result under a random folder name, and write the key to a separate file you do not open until the end. The key stays closed until every test has a written score.
The finished output is only half of the story. How a model gets there tells you a lot about how it works, and it is the most entertaining part of the comparison to watch. We had each model take periodic screenshots of its own build as it went, then stitched those into time-lapses, one per side per test.
The screenshots do real work beyond entertainment. They show whether a model builds in layers or rewrites everything, whether it checks its own work along the way, and where it got stuck. They also give you a record when something goes wrong, which matters a lot for the failure we hit in the Blender test, described below.
The setup is a line in the prompt or a small instruction asking the model to save a screenshot of its work at regular intervals into a named folder. Keep that instruction identical on both sides and keep it out of the creative part of the brief. The time-lapse is a measurement tool, so it must not change what gets built.
Quality is the first number, but it is not the only one. For each test we tracked four things: how long the run took, how many tokens it used, what it cost in dollars, and which result won. Keep them in one table with one row per test and one column per model. A plain spreadsheet works.
Pricing was identical between the two models in our run, at $2 per million input tokens and $10 per million output tokens. That made the comparison clean, because any difference in cost came from how much work each model did, and not from a different price sheet. Check the current pricing for any pair you test, because it can differ, and cached input rates can differ even when the headline numbers match.
Always record the crash and stall cases as their own line. A run that never finishes has a time and a cost, but it has no output, so its quality is zero. That is a real result, and dropping it from the table would make the model look better than it was.
Taste is the hard part, because it feels subjective. You can make it much more reliable with a short scorecard that you fill in for every result, in the same order, before you think about the winner. We judged on four questions.
Score each from one to ten, then write a single sentence explaining the lowest score. The sentence forces you to be specific, and it gives you something to compare when two results feel close. In our landing page test, one side scored a seven to nine out of ten on the strength of its scroll-driven story, and the other scored lower because the page was much shorter, with little happening in the background.
Ties are allowed, and they are useful. On our five-worlds page we called it a tie, because one side was more interactive and the other made bolder design choices. A tie tells you the two models are close on that kind of job, which is a fair finding and a reason to let price or speed decide.
Here is how our run came out. Sonnet 5.5 won five of the six tests, and the five-worlds page was the tie. The numbers below are the ones we tracked on camera, and a few are rounded as shown.
After the first three tests, Sonnet had taken about half an hour and spent about $3, and Sol had taken about 45 minutes and spent $1.27. After five tests, Sonnet had taken about 43 minutes for just under $5, and Sol 76 minutes for $2.15. The final totals were these.
One honest caveat. This is six tests, one attempt each, scored by one person. It is a strong signal about one kind of work, which is visual and interactive builds from short prompts. It says little about, for example, long refactors or data work. That is exactly why your own run, on your own jobs, is worth more than ours.
This is the idea that generalizes far past our six tests. Both models charged the same rate per token. Yet the totals were $6.68 and $2.64, and the time was 57 minutes against 94. The price sheet told us nothing about any of that, because what you pay depends on how much work the model decides to do on the job.
Sol used fewer tokens overall and spent less money, about 60 percent less. It also took about 65 percent longer, which is 37 more minutes of waiting for the same six jobs. On several tests, such as the five-worlds page, it was the slower one and still the cheaper one. On the pitch deck, Sonnet was both faster and far more expensive. The two axes do not move together.
So the useful question is what a finished result is worth to you. If you would pay a person to do the job, even a simple Blender scene can take an hour or two by hand, so an extra $1.50 for a scene that actually finishes is a bargain. If you run a thousand small background jobs where nobody is looking at the output, the cheaper run wins and the speed barely matters. Pick the model by the job and not by the price sheet.
A failed run has a hidden cost too. The crashed Blender run still spent its time and tokens, then produced nothing. When you compare cost, compare the cost of a finished, usable result, which includes the reruns you would need when a run fails.
Most bad head-to-heads fail in the same handful of ways. Knowing them in advance is most of the value, so here is the list we keep next to our own test sheet.
There is one more trap that is easy to miss. The judge can drift. By test six you are tired and you have seen a lot of polished pages, so your bar has moved. Score each test on the same day if you can, use the same scorecard every time, and read your written one-line definition of winning before each result, not after it.
This is the shortest honest path from reading this page to having a real result you trust. It takes an afternoon. Most of that time is the models working while you do something else.
Your answer will not always be 'one model wins everything'. The most useful outcome is often a split. One model for the big, high-stakes build where finishing well matters most, and another for the cheap, high-volume work where speed and price matter more.
Everything you need to run this method is on this page: what to test, how to blind it, what to measure, how to score taste, and how to read the numbers. What we keep inside the Claude Code Club is the set of exact prompts we used for these six tests, the time-lapse setup, and the scorecard in a ready-to-fill form, so you can skip the drafting and start running. Members also post their own head-to-heads there, which is the fastest way to see how other pairs of models compare on jobs like yours.
Why run the tests one-shot with no skills or references?
Because skills and references change what you are measuring. With no help, you see each model's default taste and judgment. Once you know that, you can run a second round with your real setup and see how much each model improves when it is given help.
Did Sonnet 5.5 really win almost every test?
In our run it won five of six, and the five-worlds page was a tie. That covers visual and interactive builds from short prompts, scored by one person on one attempt each. It is a strong signal for that kind of work, not a verdict on every kind of task.
If both models have the same price per token, why did the totals differ so much?
Because the bill depends on how much work each model does on a job, not on the rate. Sonnet used more tokens, about 4.5 million against about 3 million, and spent $6.68 against $2.64. Time also differed, at 57 minutes against 94, so the two models traded cost for speed in different tests.
What happened with the Blender test?
Sol's final script used a Blender setting that no longer exists, and the program crashed and did not come back. Models can often catch their own errors and fix them, but in this case it did not recover. We recorded it as a failed run, which is why the table should always include crashes.
How many tests do I need for a fair comparison?
Three is the minimum, six is comfortable. Choose jobs that test different skills, such as design, interaction and writing, so a model cannot win by being good at one trick. One test is an anecdote, and a spread of them is a pattern.
The framework is yours, and it is complete on this page. The six exact prompts, the time-lapse setup and the scorecard we use live inside the Claude Code Club, along with a room full of people running their own comparisons.
Join the Club — $9/mo