Video walkthroughs

Opus 5 vs Fable 5: The Head-To-Head Test

13 minute readUpdated June 2026Explore more

TL;DR

We gave Opus 5 and Fable 5 the same three-part brief for a fake roofing company and measured cost, time, and quality. Opus came in at $5.16 in 7 minutes 43 seconds, Fable at $5.93 in 6 minutes 20 seconds, and the quality result was a split: a draw on the landing page, Fable won the emails, Opus won the video ad. The useful takeaway is not the winner. It is that Opus does more work per task, which loses on deliverables judged by restraint and wins on deliverables judged by completeness.

The full comparison is right here on this page. Want a copy to keep or send to someone deciding between the two? Grab the standalone file.

Download the guide

Opus 5 launched at $5 per million tokens against Fable 5 at $10 per million, with benchmark scores that beat Fable across terminal coding, knowledge work, novel problem solving, and agentic search. On paper that is over before it starts. So we put both models on the same brief in two terminals, side by side, and measured what actually came out. The result was messier and far more useful than the benchmark chart, and the reason why is the whole point of this page.

Watch the full test: both models, same prompt, three deliverables

The test: one prompt, two terminals, three deliverables

Benchmarks measure a model on tasks that are easy to score. Real work is not like that, so the test has to be a real job with a real client shape. We invented one: a marketing agency has been hired by Summit Roofing, a roofing company in Phoenix, Arizona. One prompt, three deliverables, no follow-up messages allowed.

  • A complete landing page as a single HTML file. Mobile responsive, conversion focused, realistic copy, with a hero and CTA, four services, three testimonials, a contact form, and a footer.
  • A five-email cold outbound sequence targeting homeowners who need a roof replacement, with subject lines, output as a CSV with columns for email number, subject line, and body.
  • A 15-second cinematic ad with timestamps and a prompt per shot, then the actual video generated through a connected image and video service at 16 by 9 and 480p to keep the cost down.

That combination is deliberate. It spans three genuinely different skills: front end judgment, sales copywriting, and multi-step orchestration through an outside tool. A model that is good at all three is rare, and a comparison that only tests one of them tells you almost nothing. One terminal was set to Opus 5 as its default model, the other to Fable 5, and both got the identical paste.

What it cost and how long it took

Here are the two numbers everyone came for, and neither of them went the way the pricing page suggests.

  • Cost: Opus 5 spent $5.16. Fable 5 spent $5.93. Opus came in cheaper, but by 13 percent, not by the 50 percent the token price implies.
  • Time: Opus 5 took 7 minutes 43 seconds. Fable 5 took 6 minutes 20 seconds. Opus was the slower of the two by about a minute and twenty.

Read that again, because it is the most surprising result in the whole test. The model that costs half as much per token finished only slightly cheaper and took noticeably longer. Opus was doing more work to complete the same task, and it spent its price advantage doing it. Later in this guide there is a section on exactly where that extra work went, because it turns out to be the reason it won one of the three rounds.

Round 1: The landing page was a draw

Both models shipped a complete, working, mobile-responsive page. Both were good enough to hand to a client. They were not remotely the same page.

Opus went at fear. The headline reads "your roof won't survive another Phoenix monsoon." It put the contact form up in the hero on the right, so a visitor who is already convinced can book an inspection without scrolling. Underneath it stacked proof points and stats, the services, a four-step explanation of how the process works so nobody is confused about what happens next, five-star reviews, and a final call to action at the bottom.

Fable went at capability. Its headline reads "the Arizona sun is brutal on your roof. We're not." The whole page carries a bright, sunny Arizona look, with a small illustrated house graphic that gives it personality. It leads with services rather than proof, moves the proof points lower down, shows a couple of reviews, and puts a single contact form at the bottom.

There is no defensible winner here, and pretending otherwise would be dishonest. The fear angle and the capability angle are both legitimate for a roofing company, and which one converts is an empirical question that only a live test answers. The one structural difference worth noting is the form placement: a form in the hero captures the visitor who arrived ready, and a form only at the bottom asks everyone to earn their way to it. That is a conversion decision, not a taste decision, and it is the kind of thing you should be specifying in your brief rather than leaving to the model.

Round 2: Fable won the emails by writing less

The strangest moment of the test: both models produced the same subject line for email one, near enough word for word. "Quick question about your roof on" followed by the street name. Two different models, no coordination, same instinct for what a cold email subject line should do.

The bodies were not close. Opus wrote a long, technical, heavily-reasoned email. It explains that valley homes built in the late nineties and early two thousands are still on their original underlayment, that tile lasts fifty years while the felt paper under it lasts about twenty in that heat, and that the felt is the part actually keeping water out and you cannot see it from the ground. It is genuinely well argued. It is also too much for a cold email to a homeowner, too wordy, and too technical for most of the people who will receive it.

Fable wrote roughly half the length and landed harder. It notes that most roofs in the area went on fifteen to twenty years ago, that around year fifteen shingles start to dry out and crack, that things can look fine from the street, and that the free fifteen-minute inspection comes with a photo report. Then the line that wins it: "no pressure. If your roof is fine, I'll tell you it's fine." Then a single clear ask.

Fable won this round clearly, and it won it by leaving things out. Cold outbound is judged by restraint. Every extra sentence is another chance for the reader to decide this is a sales email and archive it. Opus produced the better argument and the worse email, which is a distinction worth holding onto.

Round 3: Opus won the ad by making a storyboard first

Both models wrote a 15-second script with timestamps and shot prompts, then drove a connected generation service to produce an actual video. Both videos came back watchable, which is not nothing for a one-shot brief.

Fable's ad is competent. Good push-in camera move, nice detail in the footage, a clean message about cracked tiles and curling shingles and a lifetime guarantee. It also mangles two words in the voiceover, saying "shingers" for shingles and something close to "monosu" for monsoon. That is a fixable problem, but it is a problem you have to catch.

Opus's ad is better, and the gap is not subtle. It tells an actual story across fifteen seconds: a beautiful sunny day, then the roof starts to leak, then the Summit Roofing crew arrives in a truck like they are the cavalry. It has a beginning, a turn, and a resolution, which is what separates an ad from a slideshow of nice shots.

The interesting part is why. Opus generated a full storyboard first. It produced image prompts and still images for the shots before it generated a single frame of video. That planning pass is almost certainly where the extra minute and twenty seconds went, and it is what bought the narrative coherence. Fable went more directly to the output. On a task where the deliverable is a story, planning the story first is not overhead. It is the work.

The pattern under the split: one trait, two opposite results

Line up the three rounds and something falls out that is more durable than any of the individual scores. Opus lost the emails and won the ad for exactly the same reason. It does more. It elaborates, it reasons out loud, it builds intermediate artifacts before committing to the final one.

That single trait is a liability on a deliverable judged by restraint and an asset on a deliverable judged by completeness. Which means the useful question is never "which model is better." It is "is this deliverable graded on what is present or on what is absent." Sort your work into those two buckets and the routing decision makes itself.

  • Graded on what is present: multi-step builds, anything requiring an intermediate artifact, video and campaign work, refactors, research, anything with a checklist of requirements to satisfy. More work per task is straightforwardly good here.
  • Graded on what is absent: cold email, ad copy, headlines, DMs, landing page hero copy, anything a stranger will decide about in two seconds. Extra reasoning shows up as extra words, and extra words cost you the reader.
  • The tell for the second bucket: if you can imagine yourself deleting half the output and it getting better, you are in it.

This is also why the benchmark chart and the bake-off disagreed. Benchmarks almost exclusively measure the first bucket, because completeness is easy to score automatically and restraint is not. A model can top every published chart and still write a worse cold email than the model it beat.

Price per million is not what you pay

The headline pricing difference was two to one. The measured cost difference on a real job was 13 percent. That gap is not an anomaly and it is worth internalizing before you make any budget decision off a pricing page.

Your bill is price per token multiplied by tokens consumed, and the second number is a property of the model's behavior, not of its price list. A model that plans, checks itself, and revises will consume more tokens on the identical task. Halving the token price of a model that does twice the work leaves you roughly where you started. The only number that means anything is cost per finished deliverable, and the only way to get it is to run the job.

The decision rule that outlives this comparison

Every specific number on this page has a shelf life. The next release resets the benchmarks and the pricing. The rule does not, because it is about matching a model's disposition to a deliverable's grading criterion, and that relationship holds no matter what the models are called.

Reach for the more thorough, more expensive-per-task model when:

  • The task has multiple dependent steps and a mistake in step two poisons step five.
  • The output is a build, a system, or anything that has to actually run.
  • The deliverable benefits from an intermediate artifact such as a storyboard, an outline, a schema, or a plan.
  • You will not be reviewing the output closely, so self-checking is doing your quality control for you.
  • The problem is novel enough that there is no obvious template to pattern-match against.

Reach for the faster, more direct model when:

  • The output is short-form persuasion and every extra sentence hurts.
  • You are iterating quickly and will run the task many times, where a minute of latency compounds.
  • The task is well-trodden and the model has seen ten thousand of them.
  • You are going to review and edit the output anyway, so its self-checking adds cost you will not benefit from.

You are not obliged to pick one and live with it. Both terminals were open at the same time in this test, which is the setup we would recommend as a default: run the two models side by side on anything that matters and let the work choose. It costs one extra paste.

What the published benchmarks still tell you

The bake-off is the primary evidence, but the published numbers add context a three-task test cannot reach. Opus 5 leads on terminal coding, knowledge work, novel problem solving, and agentic search. At maximum effort it lands within 5 percent of Fable 5's peak score at half the cost per task, and it more than doubles the previous Opus generation's performance at a lower cost per task. Fable at its highest effort settings still edges ahead on some measures, at significantly greater cost. Fable also remains ahead on advanced cybersecurity work such as finding vulnerabilities.

The detail worth more than any of those scores shows up in the reviews from companies who ran it in production. The recurring theme is not raw capability. It is that Opus checks and validates its own work before it calls the task done, then checks again, and it shows measurably fewer cases of doing something you did not ask for. Self-verification is the feature, and the storyboard in round three is what that looks like on a creative task.

Run this bake-off on your own work in one sitting

The whole method here is portable, and running it on work you actually do is worth more than any comparison someone else publishes. Budget an hour and about ten dollars.

  1. 1Pick one real job you would genuinely delegate. Not a puzzle, not a toy. Something a client or your own business is waiting on.
  2. 2Write it as a single brief with three deliverables that stress different skills. Include at least one that is graded on completeness and one that is graded on restraint.
  3. 3Decide your scoring criteria in writing before you run anything. What would make each deliverable good, in one sentence each. This is the step people skip and it is the step that makes the result trustworthy.
  4. 4Open two terminals, set each to a different model, and paste the identical brief into both. Change nothing else.
  5. 5Send no follow-up messages. Let each one finish or fail on its own.
  6. 6Record cost and elapsed time for each run the moment they finish, before you look at the output and your opinion contaminates the numbers.
  7. 7Score each deliverable against the criteria you wrote in step three, and note any round that is honestly a draw rather than forcing a winner.
  8. 8Write down one sentence about which kinds of task each model should get from now on. That sentence is your routing rule, and it is the actual output of the exercise.

Do this once a quarter, or whenever a major release lands. It takes an hour and it replaces a year of arguing about which model is better with an answer that is true for your work specifically.

The failure modes that make a model comparison useless

Most published comparisons are worthless for a reason. These are the ones that do the damage, and every one of them is easy to avoid once you have seen it named.

  • Testing one task type and generalizing. A single coding benchmark predicts nothing about copywriting, and the split result in this test is the proof.
  • Steering one run and not the other. The moment you give either model a correction, you are comparing your prompting, not the models.
  • Deciding the winner before you measure. The framing going into this test was that the cheaper model would dominate, and it took a clean draw and a clean loss along the way.
  • Judging by impressiveness instead of fitness. The longer, more knowledgeable email was the worse email.
  • Comparing token prices instead of finished-job cost. Two to one on the price page became 13 percent in reality.
  • Ignoring latency. A minute and twenty seconds is irrelevant on a one-off build and a serious tax on something you run fifty times a day.
  • Forcing a winner on a round that is genuinely a tie. A draw is information. Reporting it as a narrow win throws that information away.
  • Treating the result as permanent. Every number in a comparison expires at the next release. The routing rule you extracted from it does not.

The honest scoreboard from this test is one draw, one win each, cheaper and slower on one side, faster and pricier on the other. That is a less satisfying headline than "one model killed the other," and it is considerably more useful, because it tells you what to do on Monday morning with two different kinds of task in front of you.

Common questions

  • Which model actually won?

    Neither, cleanly. Across three deliverables it was a draw on the landing page, a win for Fable 5 on the cold email sequence, and a win for Opus 5 on the video ad. Opus was cheaper by 13 percent and slower by about a minute and twenty seconds. The useful result is the pattern rather than the tally: Opus does more work per task, which lost it the email round and won it the video round.

  • If Opus 5 is half the price per token, why did it only cost 13 percent less?

    Because your bill is price per token multiplied by tokens used, and Opus consumed considerably more tokens on the same job. It generated a full storyboard with image prompts before producing the video, and it checks and revises its own work as it goes. Halving the rate of a model that does more work leaves you close to where you started. Always compare cost per finished job, not the pricing page.

  • Why did the slower model produce the better video?

    It planned first. Opus generated a storyboard with still images and shot prompts before generating any video, and that planning pass is where the extra time went. The result was an ad with a real narrative arc: sunny day, roof leaks, crew arrives. On a deliverable that is fundamentally a story, planning the story before rendering it is the work, not overhead.

  • Should I just use the cheaper model for everything?

    Only if your work is mostly graded on completeness. Short-form persuasion such as cold email, ad copy, headlines, and hero copy is graded on restraint, and the more thorough model reliably overwrites there. The practical setup is to keep both available and route by deliverable type rather than committing to one.

  • How much does a test like this cost to run myself?

    This one cost $5.16 on one model and $5.93 on the other, so about eleven dollars total for three substantial deliverables built twice. Budget an hour of your time and roughly ten dollars to run the same comparison on a real job of your own, which is a far better basis for a decision than anyone else's benchmark.

  • Will this comparison still be accurate after the next model release?

    The numbers will not be. The decision rule will. Match a model's disposition to how the deliverable is graded: reach for the thorough, self-checking model on multi-step builds and anything needing an intermediate artifact, and reach for the direct, faster model on short persuasive copy and high-volume iteration. Re-run the one-hour bake-off whenever a major release lands and update your routing rule from the result.

Want the test harness and the prompts that run it?

Get 650+ plug-and-play skills, MCPs & prompts, plus 7,000+ members - $9/mo, cancel anytime.

Join the Club