Video walkthroughs

The Cinematic Intro System: Make Claude Opus 5 Direct Your Film

13 minute readUpdated June 2026Explore more

TL;DR

Do not ask an AI model for a finished video. Make it direct one. You hand it a script, and across seven separate passes it breaks the script into shots, locks your likeness into a reusable character sheet, chooses the world, draws a single composite storyboard, test renders at a quarter of the cost, critiques its own work, and only then ships the full resolution finals. The layering is the whole trick. Each pass gives the next one something solid to stand on, which is why the output looks directed instead of generated.

The full guide is right here on this page. The download is a self-contained copy you can keep, read offline, and mark up while you build.

Download the guide

Almost everyone using AI for video is doing the same thing. They describe a scene, generate a clip, look at it, generate another one, and hope the pieces cut together at the end. The results always look the same too. Technically impressive, weirdly hollow, and impossible to put in front of an audience. The problem is not the model. The problem is the job we keep handing it. A director does not generate shots. A director reads a script, decides what the story is actually about, chooses a world for it to live in, and then assigns every line a picture that carries meaning. That is a sequence of decisions, and it is exactly the sequence we are going to walk Claude Opus 5 through.

Watch the full build: giving Claude Opus 5 a script and letting it direct the entire cinematic intro end to end.

This guide covers the seven stages of that system, what decision sits inside each one, and where it tends to break. The stack is Claude Opus 5 running in Claude Code, connected to Higgsfield over MCP for image and video generation. Higgsfield matters here for one specific reason we will get to in stage three, and it is not the model quality.

Stage 1: Write the script before you think about pictures

The first pass produces spoken words and nothing else. No timestamps, no shot direction, no visual language. Ask for a thirty second intro script on your topic and explicitly forbid anything that is not dialogue. Thirty seconds is roughly seventy to ninety spoken words, which is around six to nine lines of real script.

This feels like a throwaway step and it is the one people skip most often. When you let the model write visuals and words at the same time, it writes toward whatever picture is easiest to generate. You get scripts about glowing interfaces and slow motion liquid because those render nicely. You want the script written blind, so it is written for meaning. The pictures come later and they have to serve the words, not the other way around.

There are three valid inputs at this stage and the system handles all of them. You can hand it a finished script you wrote yourself. You can hand it a full transcript of a video you already recorded and let it pull the intro out. Or you can hand it a bare topic and let it write from scratch. The third option is the fastest and the weakest, because the model has no idea what your actual video argues.

Stage 2: Break the script into a line sheet

The second pass turns the paragraph into a numbered line sheet, where every line gets exactly one image. This is the unit of work for everything that follows. One line, one shot, one picture that carries that line.

The reason to make this its own pass is that it forces a decision the model would otherwise skip. A run-on sentence with three ideas in it cannot be one picture. Splitting it makes the model commit to where the beats actually land. When it hands the sheet back it will usually flag lines that are too abstract to shoot, which is genuinely useful criticism. Lines like "innovation is accelerating" have no picture in them. Rewrite those now, while it costs you nothing.

This is also your cheapest editing window. Cutting a line here removes one shot from every downstream stage. Cutting it after the storyboard exists means regenerating work.

Stage 3: Build a character sheet once and reuse it forever

If you want to appear in the film, this is the stage that makes it possible. A character sheet is one image containing three panels of you against a seamless neutral gray studio background, generated from real reference photos. Every shot afterward references that sheet, which is how your face and wardrobe stay consistent across an entire film instead of drifting into a different person by shot four.

The inputs are simple and the quality of your inputs sets the ceiling on everything. Take about six photos of yourself against a plain white wall. Neutral expression, no smiling, one outfit, even light. That is all. Neutral gray backgrounds consistently outperform every other background for the sheet itself, so specify it.

The reason to use Higgsfield rather than any other generator is not model quality. It is asset persistence. Higgsfield saves the character sheet to your account, so you build it once and reference it in every video you make from then on. The MCP connection means Claude Code can drive all of it from your terminal, including a media upload widget that pushes your reference photos in without you touching a browser.

There are three decisions the system will ask you to make here, and they are real decisions rather than formalities:

  • Wardrobe. Whatever you pick gets locked across every panel and therefore every shot in the film. Pick something you would actually be filmed in, and pick something that survives the color grade you are about to choose in stage four.
  • Physique and age handling. Generators skew heavier and older than reality with startling consistency. You can ask for a flattering read that is still recognizably you, or a strongly idealized one. Go flattering. Idealized crosses into someone else's face, and the whole point is that the audience recognizes you.
  • Resolution. Two thousand pixel output is the sensible default. Four thousand costs more credits and buys detail that mostly matters if you plan to hold on close-ups. A low resolution sheet is the single most common cause of a face that looks almost right and unsettling.

Stage 4: Let the model choose the world before it chooses the shots

This is the pass that separates a directed film from a slideshow. Before any shot gets described, Opus reads the script for subtext, states what the video is actually promising the viewer, and then defines a single coherent world for the whole thing: the location, the time of day, the lighting, the color palette, and the camera style.

One world for the entire intro. Not a world per shot. That constraint is what makes six generated clips feel like one film, because the lighting and the grade carry across all of them even though they were rendered separately.

Once the world is set, the model works through the line sheet and assigns each line a treatment: what the camera sees, what the setting is, what the on-camera subject is doing, what supporting elements are in frame, how the camera moves, and why that image sells that specific line. That last field is the one to actually read. If the model cannot explain why a picture serves its line, the picture is decoration.

This is your last cheap intervention point. Everything after here is generation, and generation costs credits. If the world it chose is not the world you wanted, say so in plain language now. Tell it the palette is too desaturated, or that you want the whole thing warm instead of cold, and it will rewrite the treatment map for free. Fixing the same complaint after the storyboard exists costs a full regeneration.

Stage 5: Render one storyboard, not six images

The treatment map now becomes a production shot list, and then a single composite storyboard image with every panel laid out on one board.

Rendering one composite instead of six separate frames is a genuine cost decision, not a stylistic one. One image generation instead of six, for the same information. But the bigger reason is what happens next: the storyboard itself gets handed to the video model as a visual reference. It is not a planning artifact you look at and discard. It is an input. The video model sees the composition, the framing, and the grade you approved, alongside your character sheet, and generates against both.

The shot list that produces it is deliberately exhaustive. It names every asset that needs its own generation pass, and this is where a lot of people underestimate the work. Your character sheet is one asset. But if a location or a product appears in multiple shots, it needs its own reference too, or it will change between shots. A bottle in shot two and the same bottle in shot five will be two different bottles unless something anchors them.

The list also fixes lens choices, lighting, color grade, wardrobe, and the audio approach across the whole film. On audio, the default that works is diegetic sound only. The natural noise of the scene, no voiceover, no music, no spoken words in the generated clips. Your real voiceover goes over the top in the edit, and generated speech under a real voiceover is a mess nobody wants to untangle.

Stage 6: Test at 480p and let the model grade its own work

Never render finals first. Render the entire film at 480p, which costs roughly a quarter of a full resolution pass, and watch it end to end.

Everything is present at 480p except sharpness. Composition, timing, continuity, camera motion, whether your face holds, whether the cuts between the timestamped segments land. A mistake that is going to ruin the film is fully visible in the cheap version. You are paying twenty five percent to find out whether the other seventy five percent is worth spending.

The part that genuinely surprises people is that Opus 5 reviews its own output and reports what went wrong before you say anything. In the build on this page it flagged an unintended object appearing mid-shot and problems with the cuts between camera motions, both of which were easy to scroll past on a first watch. A model that verifies its own work is the difference between a tool you supervise and a tool you can hand a job to. It then asks whether you want it to fix the issues and re-test at 480p, or push straight to full resolution.

How to decide: if the problems are physical continuity, an object that should not be there, a wrong wardrobe, a face that drifted, re-test at 480p. Those failures recur. If the problems are motion smoothness or grade, go to finals, because the full resolution pass often resolves them on its own.

Stage 7: Ship the finals and stitch

The last pass is one instruction, not three. Apply the approved fixes, render every clip at full resolution, verify the output against the shot list, then stitch the clips into a single file. Batching it matters because each handoff between you and the model is a chance for context to drift, and by this point the session is carrying the script, the line sheet, the character sheet, the world, and the storyboard. Let it finish.

One session, one folder, from stage one to stage seven. Every prompt goes into the same working session so each pass can see everything that came before it. Starting a fresh chat halfway through is the quiet way to lose the world definition and wonder why the finals do not match the storyboard.

Why layering beats the perfect prompt

It is worth naming the idea underneath all seven stages, because it generalizes far past video.

The instinct with a capable model is to write one enormous prompt describing the finished thing in detail. It almost never works, and the reason is not that the model is not smart enough. It is that a single prompt forces every decision to be made simultaneously, with nothing to check them against. The model picks a world while it is picking words while it is picking camera angles, and none of those choices get to inform each other.

Break it into passes and every decision gets made against a fixed thing that already exists. The line sheet is written against a finished script. The world is chosen against a finished line sheet. The shots are written against a chosen world. The storyboard is drawn against a written shot list. Each layer stands on something solid, which is why the compounding error that ruins one-shot generation never gets started.

You can apply this to anything you build with these tools. A landing page, an ad campaign, a research report. Find the sequence of decisions a professional would actually make in order, and make the model make them in that order. The prompt quality matters far less than the sequence.

Where this breaks and what to do about it

Five failure modes account for nearly every bad run:

  • The script is too long. Thirty seconds of script is short. Nine lines is a lot. If it runs to forty five seconds you now need three clips instead of two, and the cost and the stitch complexity both jump. Time it before stage two.
  • Poor reference photos. Mixed lighting, multiple outfits, or smiling photos produce a character sheet that is nobody in particular. Six plain photos, one wall, one outfit, one expression.
  • Abstract lines that have no picture. If a line cannot be photographed, the model will invent generic filler for it. Rewrite the line at stage two, do not try to fix it with a better shot description at stage four.
  • Unanchored recurring objects. Any product or location appearing in more than one shot needs its own generated reference, exactly like your face does. Otherwise it silently changes between shots.
  • Skipping the 480p test. Every single time this gets skipped to save a few minutes, it costs a full resolution regeneration instead. The test is not optional, it is the cheapest insurance in the system.

A worked example: orange soda

It helps to see the whole chain run on something deliberately mundane, because a mundane topic proves the system rather than the subject matter.

The topic was orange soda. Stage one produced a script opening on a line about a specific shade of orange that does not exist in nature, one that stains your tongue and tastes less like fruit than like a summer afternoon someone bottled by accident. Nothing visual in the prompt, just words that mean something.

Stage four is where it got interesting. Opus read the script and named the actual promise: the video will tell you what orange soda really is, its chemistry, its inventor, and why a century of bans, lawsuits, and reformulations never replaced it. Then it chose a world nobody would have thought to ask for. A roadside American gas station across one bleached afternoon. Cold fluorescent light, bleached concrete, everything desaturated except the soda itself.

That last detail is a real directorial choice. Draining the color from the entire frame so the one orange object carries all of it is the kind of decision that makes an intro look designed. It came out of stage four for free, because stage four's only job was to make that decision and it had a finished script to make it against.

From there: a composite storyboard including a hand closing on a fizzing bottle in extreme close-up and a walk down an aisle stacked with the stuff, a 480p test that caught an object appearing where it should not have plus rough cuts between camera motions, fixes applied, and a stitched 1080p final. The same chain was run on candy, quantum computing, and a golf ball on Mars without changing anything but the topic.

Your first run, start to finish

The shortest honest path from finishing this page to having a real film, in one sitting:

  1. 1Take six reference photos against a plain white wall right now. Neutral face, one outfit, even light. This is the only part that needs a camera and it takes two minutes.
  2. 2Connect Higgsfield to Claude Code as a custom MCP connector, using the MCP URL from the Higgsfield MCP and CLI page. Sign in when prompted.
  3. 3Create an empty folder for this film and open a Claude Code session in it. Select Opus 5 and set reasoning effort to high. One folder, one session, all seven stages.
  4. 4Write the thirty second script first, words only, and read it against a timer before you go further.
  5. 5Run the line sheet pass and rewrite any line the model flags as too abstract to shoot.
  6. 6Build the character sheet at two thousand pixels. Choose a wardrobe you can live with and a flattering rather than idealized physique.
  7. 7Run the world and treatment pass, then actually read the world description. Push back in plain language if it is not what you wanted. This is your free edit.
  8. 8Generate the composite storyboard and check that every recurring object has its own reference asset.
  9. 9Render the whole film at 480p and watch it end to end. Read the model's own critique before you write yours.
  10. 10Apply fixes, ship the 1080p finals, and stitch. Record your real voiceover over the top in your editor.

The first run takes an afternoon because you are learning the shape of it. The second run takes about as long as it used to take to find usable stock footage, and the character sheet is already sitting in your account waiting.

Common questions

  • Do I have to appear in the film?

    No. The character sheet stage exists so that you can, but a film with no on-camera subject skips stage three entirely and runs the other six stages unchanged. Everything else in the system is about the script, the world, and the shots.

  • Why Higgsfield instead of another image and video tool?

    Asset persistence. Higgsfield saves your character sheet and other reference assets to your account, so you build your likeness once and reference it in every film you make afterward. It also exposes an MCP connection, which means Claude Code can drive the whole generation pipeline from your terminal instead of you clicking through a browser.

  • How much does one thirty second intro cost to make?

    It depends on your resolutions and how many test passes you run, so treat the structure rather than a number as the answer. The two levers that dominate everything else are rendering one composite storyboard instead of separate frames, and testing the entire film at 480p for roughly a quarter of a full resolution pass before you commit.

  • Can I use a script I already wrote, or a transcript of a video I recorded?

    Yes to both, and both are better inputs than a bare topic. Hand it a finished script and it goes straight to the line sheet. Hand it a full transcript and it will pull the intro out of your actual video, which means the intro promises something the video really delivers.

  • What if the world it picks is not what I wanted?

    Tell it, in plain language, before you generate anything. Stage four is a text pass, so rewriting the world costs nothing. Saying the palette is too cold or that you want everything bright and warm will get you a rewritten treatment map for free. The expensive version of this mistake is trying to correct a world-level problem shot by shot after the storyboard exists.

  • Should I generate the voiceover too?

    No. Keep the generated clips to diegetic sound only, meaning the natural noise of the scene with no speech and no music, and lay your real voiceover over the top in your editor. Generated speech underneath a real voiceover is difficult to remove and adds nothing you cannot do better in the edit.

Want the skill that runs all seven stages?

Get 650+ plug-and-play skills, MCPs & prompts, plus 7,000+ members - $9/mo, cancel anytime.

Join the Club