How to Run Your Own Blind Model Head-to-Head (Sonnet 5.5 vs GPT-6.1 Sol)
TL;DR
Run two models on the same six one-shot jobs, hide which side is which, and score the output before you look at the bill. In our run Sonnet 5.5 won five of six tests and finished in 57 minutes against 94 for GPT-6.1 Sol, while Sol cost $2.64 against $6.68. Cost per task is not cost per token, so pick the model by the job and not by the price sheet.
The full guide is right here on this page. The download is the same guide as a single offline file you can keep.
Download the guideMost model comparisons you read online are one person's vibe after a few chats. A better way exists, and it fits in an afternoon. You give two models the same short brief, you hide which one made which result, you pick the winner with your own eyes, and only then do you look at the clock and the bill. This guide walks through that method and uses our own six-test run of Claude Sonnet 5.5 against GPT-6.1 Sol as the worked example.
Step 1: Decide what question you are actually asking
A head-to-head without a question produces a pile of screenshots and no decision. Before you run anything, write one sentence that starts with 'I am choosing a model for'. For us it was 'building websites, games and other visual things from a short prompt with no scaffolding'. That sentence decides which tests are worth running and which are a waste of an hour.
There are two honest kinds of question. The first is 'which model should be my default', which needs a spread of jobs. The second is 'which model should run this one recurring task', which needs one test repeated a few times. This guide covers the first kind, because it is the harder one to get right, and the second kind is just a smaller version of the same process.
Write the question down where you can see it while you score. The question is the judge, and your mood on the day is not. When a result is impressive but has nothing to do with your question, it does not count.
Step 2: Pick use cases that expose differences
Easy tasks make every model look the same. If both models can do the job in their sleep, you learn nothing. The tests worth running are the ones where the ceiling is high and the results can be compared by eye in under a minute. We used six:
- 1A landing page that stops the scroll, for a brand the model has to invent.
- 2A 3D browser game, playable from start screen to game over.
- 3A striking 3D scene built in Blender entirely from code.
- 4One page holding five fundamentally different visual worlds.
- 5An interactive 3D explainer of something hard to picture, like how a black hole bends light.
- 6A ten-slide investor pitch deck that runs in the browser.
Look at what those six have in common. Every one of them ends in something you can open and judge immediately. Every one has a wide range of possible quality, so a gap between models shows up instead of hiding. And every one mixes several skills at once, such as design, code, interaction and writing, so a model has to be good at the whole job and not at one trick.
Each brief was short and open. Each one told the model what to make and what a good result feels like, then left every choice to the model. Brand, genre, subject, palette and structure were all up to the model, which is exactly why the results separate. The exact wording of our six prompts lives inside the Claude Code Club, and the decision to keep them short is the part you can copy today.
Step 3: Write minimal prompts and strip out the help
We gave both models very little. No skills, no reference images, no design system, no example code, and no follow-up turns. One prompt each, one attempt each. That is what we mean by one-shot, and it is the most important rule in the whole method.
The reason is simple. The moment you add skills, references or coaching, you are measuring how well each model uses your help, and the help itself becomes a hidden variable. If you want to know which model has better taste and judgment by default, you have to take the scaffolding away. You can test a full setup later as a second round.
The other half of the rule is symmetry. Both models get the identical text, the same tools and the same empty starting folder. If one model needs a different harness or an extra instruction to even start, write that down as a finding, because it counts against ease of use. Do not quietly fix it for one side.
- Same prompt text, pasted once, for both models.
- Same starting folder, with nothing in it.
- No skills, references or example files for either side.
- One attempt per test. No regenerating until it looks good.
- If a model crashes or stalls, record that and do not rescue it.
Step 4: Run it blind so your taste is not biased
People are bad at judging work once they know who made it. If you believe one model is better, you will forgive its rough edges and notice the other one's. Blinding fixes this cheaply. Have each model write its output into a folder labeled only A or B, let a coin flip or a script decide which model goes in which folder, and keep the key somewhere you will not look until you have scored all six tests.
In our run we did not know which side was which until we had picked a winner for each test. On the landing page test, for example, the left side was clearly richer, with parallax, a mouse-following glow, and a page that moved from morning to night as you scrolled. We picked it before seeing whose it was. Only then did the reveal show it was Sonnet 5.5, and the other side was Sol.
If you are testing alone, the simplest blind setup is to ask a second Claude Code session to set it up for you. Tell it to run both models, save each result under a random folder name, and write the key to a separate file you do not open until the end. The key stays closed until every test has a written score.
Step 5: Capture time-lapses so you can see the process
The finished output is only half of the story. How a model gets there tells you a lot about how it works, and it is the most entertaining part of the comparison to watch. We had each model take periodic screenshots of its own build as it went, then stitched those into time-lapses, one per side per test.
The screenshots do real work beyond entertainment. They show whether a model builds in layers or rewrites everything, whether it checks its own work along the way, and where it got stuck. They also give you a record when something goes wrong, which matters a lot for the failure we hit in the Blender test, described below.
The setup is a line in the prompt or a small instruction asking the model to save a screenshot of its work at regular intervals into a named folder. Keep that instruction identical on both sides and keep it out of the creative part of the brief. The time-lapse is a measurement tool, so it must not change what gets built.
Step 6: Measure time, tokens and cost for every test
Quality is the first number, but it is not the only one. For each test we tracked four things: how long the run took, how many tokens it used, what it cost in dollars, and which result won. Keep them in one table with one row per test and one column per model. A plain spreadsheet works.
Pricing was identical between the two models in our run, at $2 per million input tokens and $10 per million output tokens. That made the comparison clean, because any difference in cost came from how much work each model did, and not from a different price sheet. Check the current pricing for any pair you test, because it can differ, and cached input rates can differ even when the headline numbers match.
- Time is wall-clock minutes from sending the prompt to a finished, working result.
- Tokens is total usage for the run, input plus output, read from your usage screen.
- Cost is the dollar figure for that run, which is tokens times the rate.
- Quality is your blind score out of ten, plus a one-line reason.
Always record the crash and stall cases as their own line. A run that never finishes has a time and a cost, but it has no output, so its quality is zero. That is a real result, and dropping it from the table would make the model look better than it was.
Step 7: Score taste with a scorecard, not a feeling
Taste is the hard part, because it feels subjective. You can make it much more reliable with a short scorecard that you fill in for every result, in the same order, before you think about the winner. We judged on four questions.
- Does it work? Every button, scroll and interaction does what it says, with no dead ends.
- Does it look designed? Clear hierarchy, consistent spacing, intentional color and type, without the generic tells of default AI output.
- Is there depth? Hover states, motion, hidden interactions and details that reward a second look.
- Would you ship or show it? The gut check at the end, after the first three.
Score each from one to ten, then write a single sentence explaining the lowest score. The sentence forces you to be specific, and it gives you something to compare when two results feel close. In our landing page test, one side scored a seven to nine out of ten on the strength of its scroll-driven story, and the other scored lower because the page was much shorter, with little happening in the background.
Ties are allowed, and they are useful. On our five-worlds page we called it a tie, because one side was more interactive and the other made bolder design choices. A tie tells you the two models are close on that kind of job, which is a fair finding and a reason to let price or speed decide.
Step 8: The worked example, six tests and the results
Here is how our run came out. Sonnet 5.5 won five of the six tests, and the five-worlds page was the tie. The numbers below are the ones we tracked on camera, and a few are rounded as shown.
- Landing page: Sonnet 5.5 won by a wide margin. It took 7 minutes 44 seconds and cost $0.91. Sol took about ten minutes and cost roughly a third of that, but the page was shorter and flatter.
- 3D browser game: Sonnet 5.5 won. It built a neon arcade game that we scored a ten out of ten, in about 9.5 minutes. Sol took about 17.5 minutes and produced a playable driving game that scored a five or six.
- Blender scene: Sonnet 5.5 won, and this was the failure test. Sol's final script used a Blender setting that no longer exists, crashed Blender, and never recovered. The Sonnet run produced a full scene with geometry, materials, lights, water and sky.
- Five visual worlds: a tie. Sonnet took about 7 minutes and cost $0.80, while Sol took about 18 minutes and cost $0.50.
- Interactive 3D explainer: Sonnet 5.5 won with a black hole light-bending simulation, in about 8 minutes for roughly $1. Sol built a four-stroke engine explainer in 13 minutes for $0.37.
- Investor pitch deck: Sonnet 5.5 won, in 13 minutes against 17. Its deck cost about four times as much, and it had the better layout, animation and team slide.
After the first three tests, Sonnet had taken about half an hour and spent about $3, and Sol had taken about 45 minutes and spent $1.27. After five tests, Sonnet had taken about 43 minutes for just under $5, and Sol 76 minutes for $2.15. The final totals were these.
- Sonnet 5.5: 57 minutes, just over 4.5 million tokens, $6.68.
- GPT-6.1 Sol: 94 minutes, just over 3 million tokens, $2.64.
One honest caveat. This is six tests, one attempt each, scored by one person. It is a strong signal about one kind of work, which is visual and interactive builds from short prompts. It says little about, for example, long refactors or data work. That is exactly why your own run, on your own jobs, is worth more than ours.
Step 9: Cost per task is not cost per token
This is the idea that generalizes far past our six tests. Both models charged the same rate per token. Yet the totals were $6.68 and $2.64, and the time was 57 minutes against 94. The price sheet told us nothing about any of that, because what you pay depends on how much work the model decides to do on the job.
Sol used fewer tokens overall and spent less money, about 60 percent less. It also took about 65 percent longer, which is 37 more minutes of waiting for the same six jobs. On several tests, such as the five-worlds page, it was the slower one and still the cheaper one. On the pitch deck, Sonnet was both faster and far more expensive. The two axes do not move together.
So the useful question is what a finished result is worth to you. If you would pay a person to do the job, even a simple Blender scene can take an hour or two by hand, so an extra $1.50 for a scene that actually finishes is a bargain. If you run a thousand small background jobs where nobody is looking at the output, the cheaper run wins and the speed barely matters. Pick the model by the job and not by the price sheet.
A failed run has a hidden cost too. The crashed Blender run still spent its time and tokens, then produced nothing. When you compare cost, compare the cost of a finished, usable result, which includes the reruns you would need when a run fails.
Step 10: Avoid the failure modes that ruin a comparison
Most bad head-to-heads fail in the same handful of ways. Knowing them in advance is most of the value, so here is the list we keep next to our own test sheet.
- Unblinded scoring. You already know which model you like, so you score it higher. Blind it.
- Regenerating until it looks good. One attempt per test, or you are measuring your patience.
- Helping one side. A skill, a reference or a nudge for only one model breaks the comparison.
- Trusting one test. One win is an anecdote. Six tests across different skills is a pattern.
- Ignoring the crash. A run that fails outright belongs in the table as a zero.
- Comparing tokens, not cost. Different models and different cache rules can make equal token counts cost different amounts.
There is one more trap that is easy to miss. The judge can drift. By test six you are tired and you have seen a lot of polished pages, so your bar has moved. Score each test on the same day if you can, use the same scorecard every time, and read your written one-line definition of winning before each result, not after it.
Step 11: Run your own head-to-head in one sitting
This is the shortest honest path from reading this page to having a real result you trust. It takes an afternoon. Most of that time is the models working while you do something else.
- 1Write your one-sentence question, starting with 'I am choosing a model for'.
- 2Pick three to six use cases that end in something you can open and judge in a minute, and include one you expect the other model to win.
- 3Write one short, open prompt per use case. Say what to make and what good feels like, and leave every choice to the model.
- 4Set up two empty folders, label them A and B, and let a coin flip or script decide which model goes where. Keep the key in a file you do not open.
- 5Add an identical instruction to both runs to save a screenshot at regular intervals, so you get time-lapses.
- 6Run each test once per model with no skills, references or coaching. Record time, tokens and cost as each one finishes.
- 7Score every result with the four-question scorecard before you open the key. Write one sentence for each lowest score.
- 8Open the key, add up the wins, and compare time and cost per finished result. Decide which model runs which kind of job.
Your answer will not always be 'one model wins everything'. The most useful outcome is often a split. One model for the big, high-stakes build where finishing well matters most, and another for the cheap, high-volume work where speed and price matter more.
Step 12: Where the exact prompts and templates live
Everything you need to run this method is on this page: what to test, how to blind it, what to measure, how to score taste, and how to read the numbers. What we keep inside the Claude Code Club is the set of exact prompts we used for these six tests, the time-lapse setup, and the scorecard in a ready-to-fill form, so you can skip the drafting and start running. Members also post their own head-to-heads there, which is the fastest way to see how other pairs of models compare on jobs like yours.
Common questions
Why run the tests one-shot with no skills or references?
Because skills and references change what you are measuring. With no help, you see each model's default taste and judgment. Once you know that, you can run a second round with your real setup and see how much each model improves when it is given help.
Did Sonnet 5.5 really win almost every test?
In our run it won five of six, and the five-worlds page was a tie. That covers visual and interactive builds from short prompts, scored by one person on one attempt each. It is a strong signal for that kind of work, not a verdict on every kind of task.
If both models have the same price per token, why did the totals differ so much?
Because the bill depends on how much work each model does on a job, not on the rate. Sonnet used more tokens, about 4.5 million against about 3 million, and spent $6.68 against $2.64. Time also differed, at 57 minutes against 94, so the two models traded cost for speed in different tests.
What happened with the Blender test?
Sol's final script used a Blender setting that no longer exists, and the program crashed and did not come back. Models can often catch their own errors and fix them, but in this case it did not recover. We recorded it as a failed run, which is why the table should always include crashes.
How many tests do I need for a fair comparison?
Three is the minimum, six is comfortable. Choose jobs that test different skills, such as design, interaction and writing, so a model cannot win by being good at one trick. One test is an anecdote, and a spread of them is a pattern.
Keep going
- Opus 5 vs Fable 5: The Head-To-Head TestVideo walkthroughs · 13 min
- Claude Opus 5.5: What Changed and How to Use ItVideo walkthroughs · 12 min
- The Closed Loop: How to Make an AI Judge Its Own Visuals and Fix ThemVideo walkthroughs · 13 min
- How To Build A $10,000 Mobile Microsite With Claude CodeVideo walkthroughs · 14 min
Want the exact test prompts and the judging template?
Get 650+ plug-and-play skills, MCPs & prompts, plus 8,000+ members - $9/mo, cancel anytime.
Join the Club