Video walkthroughs · Claude Code Club

Opus 5 vs Fable 5: The Head-To-Head Test

13 minute read · Companion guide to “Claude Opus 5 vs Fable 5 (Head To Head)”

TL;DR

We gave Opus 5 and Fable 5 the same three-part brief for a fake roofing company and measured cost, time, and quality. Opus came in at $5.16 in 7 minutes 43 seconds, Fable at $5.93 in 6 minutes 20 seconds, and the quality result was a split: a draw on the landing page, Fable won the emails, Opus won the video ad. The useful takeaway is not the winner. It is that Opus does more work per task, which loses on deliverables judged by restraint and wins on deliverables judged by completeness.

Opus 5 launched at $5 per million tokens against Fable 5 at $10 per million, with benchmark scores that beat Fable across terminal coding, knowledge work, novel problem solving, and agentic search. On paper that is over before it starts. So we put both models on the same brief in two terminals, side by side, and measured what actually came out. The result was messier and far more useful than the benchmark chart, and the reason why is the whole point of this page.

Watch the full test: both models, same prompt, three deliverables

Watch the full test: both models, same prompt, three deliverables

The test: one prompt, two terminals, three deliverables

Benchmarks measure a model on tasks that are easy to score. Real work is not like that, so the test has to be a real job with a real client shape. We invented one: a marketing agency has been hired by Summit Roofing, a roofing company in Phoenix, Arizona. One prompt, three deliverables, no follow-up messages allowed.

That combination is deliberate. It spans three genuinely different skills: front end judgment, sales copywriting, and multi-step orchestration through an outside tool. A model that is good at all three is rare, and a comparison that only tests one of them tells you almost nothing. One terminal was set to Opus 5 as its default model, the other to Fable 5, and both got the identical paste.

What it cost and how long it took

Here are the two numbers everyone came for, and neither of them went the way the pricing page suggests.

Read that again, because it is the most surprising result in the whole test. The model that costs half as much per token finished only slightly cheaper and took noticeably longer. Opus was doing more work to complete the same task, and it spent its price advantage doing it. Later in this guide there is a section on exactly where that extra work went, because it turns out to be the reason it won one of the three rounds.

Round 1: The landing page was a draw

Both models shipped a complete, working, mobile-responsive page. Both were good enough to hand to a client. They were not remotely the same page.

Opus went at fear. The headline reads "your roof won't survive another Phoenix monsoon." It put the contact form up in the hero on the right, so a visitor who is already convinced can book an inspection without scrolling. Underneath it stacked proof points and stats, the services, a four-step explanation of how the process works so nobody is confused about what happens next, five-star reviews, and a final call to action at the bottom.

Fable went at capability. Its headline reads "the Arizona sun is brutal on your roof. We're not." The whole page carries a bright, sunny Arizona look, with a small illustrated house graphic that gives it personality. It leads with services rather than proof, moves the proof points lower down, shows a couple of reviews, and puts a single contact form at the bottom.

There is no defensible winner here, and pretending otherwise would be dishonest. The fear angle and the capability angle are both legitimate for a roofing company, and which one converts is an empirical question that only a live test answers. The one structural difference worth noting is the form placement: a form in the hero captures the visitor who arrived ready, and a form only at the bottom asks everyone to earn their way to it. That is a conversion decision, not a taste decision, and it is the kind of thing you should be specifying in your brief rather than leaving to the model.

Round 2: Fable won the emails by writing less

The strangest moment of the test: both models produced the same subject line for email one, near enough word for word. "Quick question about your roof on" followed by the street name. Two different models, no coordination, same instinct for what a cold email subject line should do.

The bodies were not close. Opus wrote a long, technical, heavily-reasoned email. It explains that valley homes built in the late nineties and early two thousands are still on their original underlayment, that tile lasts fifty years while the felt paper under it lasts about twenty in that heat, and that the felt is the part actually keeping water out and you cannot see it from the ground. It is genuinely well argued. It is also too much for a cold email to a homeowner, too wordy, and too technical for most of the people who will receive it.

Fable wrote roughly half the length and landed harder. It notes that most roofs in the area went on fifteen to twenty years ago, that around year fifteen shingles start to dry out and crack, that things can look fine from the street, and that the free fifteen-minute inspection comes with a photo report. Then the line that wins it: "no pressure. If your roof is fine, I'll tell you it's fine." Then a single clear ask.

Fable won this round clearly, and it won it by leaving things out. Cold outbound is judged by restraint. Every extra sentence is another chance for the reader to decide this is a sales email and archive it. Opus produced the better argument and the worse email, which is a distinction worth holding onto.

Round 3: Opus won the ad by making a storyboard first

Both models wrote a 15-second script with timestamps and shot prompts, then drove a connected generation service to produce an actual video. Both videos came back watchable, which is not nothing for a one-shot brief.

Fable's ad is competent. Good push-in camera move, nice detail in the footage, a clean message about cracked tiles and curling shingles and a lifetime guarantee. It also mangles two words in the voiceover, saying "shingers" for shingles and something close to "monosu" for monsoon. That is a fixable problem, but it is a problem you have to catch.

Opus's ad is better, and the gap is not subtle. It tells an actual story across fifteen seconds: a beautiful sunny day, then the roof starts to leak, then the Summit Roofing crew arrives in a truck like they are the cavalry. It has a beginning, a turn, and a resolution, which is what separates an ad from a slideshow of nice shots.

The interesting part is why. Opus generated a full storyboard first. It produced image prompts and still images for the shots before it generated a single frame of video. That planning pass is almost certainly where the extra minute and twenty seconds went, and it is what bought the narrative coherence. Fable went more directly to the output. On a task where the deliverable is a story, planning the story first is not overhead. It is the work.

The pattern under the split: one trait, two opposite results

Line up the three rounds and something falls out that is more durable than any of the individual scores. Opus lost the emails and won the ad for exactly the same reason. It does more. It elaborates, it reasons out loud, it builds intermediate artifacts before committing to the final one.

That single trait is a liability on a deliverable judged by restraint and an asset on a deliverable judged by completeness. Which means the useful question is never "which model is better." It is "is this deliverable graded on what is present or on what is absent." Sort your work into those two buckets and the routing decision makes itself.

This is also why the benchmark chart and the bake-off disagreed. Benchmarks almost exclusively measure the first bucket, because completeness is easy to score automatically and restraint is not. A model can top every published chart and still write a worse cold email than the model it beat.

Price per million is not what you pay

The headline pricing difference was two to one. The measured cost difference on a real job was 13 percent. That gap is not an anomaly and it is worth internalizing before you make any budget decision off a pricing page.

Your bill is price per token multiplied by tokens consumed, and the second number is a property of the model's behavior, not of its price list. A model that plans, checks itself, and revises will consume more tokens on the identical task. Halving the token price of a model that does twice the work leaves you roughly where you started. The only number that means anything is cost per finished deliverable, and the only way to get it is to run the job.

The decision rule that outlives this comparison

Every specific number on this page has a shelf life. The next release resets the benchmarks and the pricing. The rule does not, because it is about matching a model's disposition to a deliverable's grading criterion, and that relationship holds no matter what the models are called.

Reach for the more thorough, more expensive-per-task model when:

Reach for the faster, more direct model when:

You are not obliged to pick one and live with it. Both terminals were open at the same time in this test, which is the setup we would recommend as a default: run the two models side by side on anything that matters and let the work choose. It costs one extra paste.

What the published benchmarks still tell you

The bake-off is the primary evidence, but the published numbers add context a three-task test cannot reach. Opus 5 leads on terminal coding, knowledge work, novel problem solving, and agentic search. At maximum effort it lands within 5 percent of Fable 5's peak score at half the cost per task, and it more than doubles the previous Opus generation's performance at a lower cost per task. Fable at its highest effort settings still edges ahead on some measures, at significantly greater cost. Fable also remains ahead on advanced cybersecurity work such as finding vulnerabilities.

The detail worth more than any of those scores shows up in the reviews from companies who ran it in production. The recurring theme is not raw capability. It is that Opus checks and validates its own work before it calls the task done, then checks again, and it shows measurably fewer cases of doing something you did not ask for. Self-verification is the feature, and the storyboard in round three is what that looks like on a creative task.

Run this bake-off on your own work in one sitting

The whole method here is portable, and running it on work you actually do is worth more than any comparison someone else publishes. Budget an hour and about ten dollars.

  1. Pick one real job you would genuinely delegate. Not a puzzle, not a toy. Something a client or your own business is waiting on.
  2. Write it as a single brief with three deliverables that stress different skills. Include at least one that is graded on completeness and one that is graded on restraint.
  3. Decide your scoring criteria in writing before you run anything. What would make each deliverable good, in one sentence each. This is the step people skip and it is the step that makes the result trustworthy.
  4. Open two terminals, set each to a different model, and paste the identical brief into both. Change nothing else.
  5. Send no follow-up messages. Let each one finish or fail on its own.
  6. Record cost and elapsed time for each run the moment they finish, before you look at the output and your opinion contaminates the numbers.
  7. Score each deliverable against the criteria you wrote in step three, and note any round that is honestly a draw rather than forcing a winner.
  8. Write down one sentence about which kinds of task each model should get from now on. That sentence is your routing rule, and it is the actual output of the exercise.

Do this once a quarter, or whenever a major release lands. It takes an hour and it replaces a year of arguing about which model is better with an answer that is true for your work specifically.

The failure modes that make a model comparison useless

Most published comparisons are worthless for a reason. These are the ones that do the damage, and every one of them is easy to avoid once you have seen it named.

The honest scoreboard from this test is one draw, one win each, cheaper and slower on one side, faster and pricier on the other. That is a less satisfying headline than "one model killed the other," and it is considerably more useful, because it tells you what to do on Monday morning with two different kinds of task in front of you.

Common questions

Which model actually won?

Neither, cleanly. Across three deliverables it was a draw on the landing page, a win for Fable 5 on the cold email sequence, and a win for Opus 5 on the video ad. Opus was cheaper by 13 percent and slower by about a minute and twenty seconds. The useful result is the pattern rather than the tally: Opus does more work per task, which lost it the email round and won it the video round.

If Opus 5 is half the price per token, why did it only cost 13 percent less?

Because your bill is price per token multiplied by tokens used, and Opus consumed considerably more tokens on the same job. It generated a full storyboard with image prompts before producing the video, and it checks and revises its own work as it goes. Halving the rate of a model that does more work leaves you close to where you started. Always compare cost per finished job, not the pricing page.

Why did the slower model produce the better video?

It planned first. Opus generated a storyboard with still images and shot prompts before generating any video, and that planning pass is where the extra time went. The result was an ad with a real narrative arc: sunny day, roof leaks, crew arrives. On a deliverable that is fundamentally a story, planning the story before rendering it is the work, not overhead.

Should I just use the cheaper model for everything?

Only if your work is mostly graded on completeness. Short-form persuasion such as cold email, ad copy, headlines, and hero copy is graded on restraint, and the more thorough model reliably overwrites there. The practical setup is to keep both available and route by deliverable type rather than committing to one.

How much does a test like this cost to run myself?

This one cost $5.16 on one model and $5.93 on the other, so about eleven dollars total for three substantial deliverables built twice. Budget an hour of your time and roughly ten dollars to run the same comparison on a real job of your own, which is a far better basis for a decision than anyone else's benchmark.

Will this comparison still be accurate after the next model release?

The numbers will not be. The decision rule will. Match a model's disposition to how the deliverable is graded: reach for the thorough, self-checking model on multi-step builds and anything needing an intermediate artifact, and reach for the direct, faster model on short persuasive copy and high-volume iteration. Re-run the one-hour bake-off whenever a major release lands and update your routing rule from the result.

Want the test harness and the prompts that run it?

The method above is yours to run by hand today. The exact multi-deliverable test prompt, the scoring sheet, and the routing setup that switches models per task type live inside the club, with 7,000+ members running the same builds. $9/mo, cancel anytime.

Join the Club — $9/mo

Read this online at claudecodeclub.ai