Video walkthroughs · Claude Code Club

Build the Cast Before You Shoot the Scene

10 minute read

TL;DR

Almost everyone writes the scene first and then wonders why the character changes face between cuts. The fix is to build a cast and a set before you shoot: neutral character and product sheets on flat grey, environment plates carrying the film's actual light, all saved as named Elements you can tag. Then the scene prompt is mostly pointing.

Seedance 2.5 on Higgsfield will render a beautiful five second shot from a sentence. That is not the hard part. The hard part is the second shot, and the fifth, where the same person has to walk into the same room wearing the same jacket. Almost everyone hits that wall and tries to solve it with a longer scene prompt. It does not work, because the problem is not the scene. The problem is that there was never a cast.

Stage 1: Build the cast before you write a single scene

A scene prompt is mostly pointing. It says who is in the shot, where they are, what happens, and how it looks. Three of those four point at something that has to already exist. If nothing exists, the model invents a person on the spot, and invents a slightly different one the next time you ask. Consistency is not a prompt trick. It is a consequence of having something fixed to refer to.

So the order is fixed. Characters, products, and locations get built first, as their own generations, with their own rules. If you have a scene idea and nothing built, you are not ready to write the prompt yet. The asset stage takes twenty minutes and it is the whole difference between a film and a pile of clips that share a topic.

Stage 2: Learn the one law that governs everything else

Here is the sentence the whole workflow hangs on. Assets are neutral. Environments carry the look. Seedance marries them.

A character sheet is a spec, not a shot. Its job is to describe a person precisely enough that the model can rebuild them anywhere, so it gets flat shadowless light on a plain mid-grey field: no colour cast, no shadow direction, no mood. An environment plate is the exact opposite. Its job is to supply light, so it is generated in the film's real look, with the real lamps and the real time of day. Seedance then relights the neutral person into the lit space.

Stage 3: Use mid-grey, not white and not black

The background value on an asset sheet is a real decision, not a default. Three options exist and only one holds up.

Specify it as a hex value, not the word grey, because the word drifts across generations and the number does not. It is no coincidence that every cyclorama in the world is painted roughly this value.

textplain solid #7f7f7f mid-grey seamless studio background, flat even
shadow-free studio lighting, matte surfaces with no sheen, 85mm lens,
f/8, photorealistic

Stage 4: Put four panels on one sheet, and stop at four

A character is one generation, not twelve. One image at 2K, four panels: full-body front, full-body three-quarter, full-body back, and a tight face close-up. The three-quarter is the most load-bearing panel, because it is the only view carrying depth and silhouette at the same time, and it is what the model leans on hardest when placing a body in space.

Four is the ceiling, not a suggestion. Add a fifth and every panel shrinks inside the same canvas. The face softens first, and a soft face is exactly what you built the sheet to prevent. If a scene needs a view the sheet does not have, generate that one view later as its own image, referencing the sheet.

Two things decide whether a sheet works. State the panel count in the first sentence, or the generator quietly returns three panels. And say what each panel is responsible for showing, or the back view comes back generic because nothing told it what it was for.

Stage 5: Leave exactly one readable face on the sheet

On the full-body front panel, turn the head slightly down and away so the face there is not clearly legible. The close-up panel should be the only face on the sheet the model can read. This sounds wrong until you have watched a character drift.

The reason is averaging. Two renderings of one face on a single sheet, one large and soft and one small and sharp, are two competing sources of truth. The model splits the difference, and it splits it slightly differently every time. Give it one source of truth and the face stops moving. Skip this and the drift shows up around shot five, late enough that you have already built on top of it.

Here is the full prompt for that. Fill in the short list at the top and leave blank anything you do not know. The rule at the top tells the model to work the rest out and to tell you what it decided, so a blank is a choice you are handing over, not a gap.

textFILL-IN RULE: Fill in the short list below. Leave blank anything you do not
know. Then write the finished prompt: take my answers and fold them into the
template underneath, so what comes back is one clean prompt with no brackets
left in it and no leftover list on top. Anything I left blank, work it out
from the rest of the prompt yourself. Do not ask me first. At the end, tell me
in two lines what you decided on my behalf so I can change anything you got
wrong.

AGE:
BUILD (height in cm, shoulders, how they stand):
HAIR (length, colour, how it sits, parting, finish):
FACE (brows, eyes, nose, face shape, facial hair):
DISTINGUISHING MARKS (one or two):
CLOTHING (each garment: fabric, cut, colour, how it hangs, condition):
FOOTWEAR:

--- THE PROMPT (leave this part alone) ---

A four-panel character sheet. This image contains four views in total: full-body
front, full-body three-quarter, full-body back, and a tight head-and-shoulders
close-up. All four must be present and all four must show the same person.

Panel 1 - full-body front, standing straight and facing the camera in a relaxed
pose, arms hanging at their sides, full height from head to toe, no cropping at
the feet or the top of the head. Head tilted slightly down and away so the face
is not clearly readable.
Panel 2 - full-body three-quarter, rotated 45 degrees to their left so both the
front and one side of the body are visible.
Panel 3 - full-body back, seen from directly behind, face not visible at all.
Panel 4 - tight head-and-shoulders close-up, front, looking directly into the
lens, mouth closed, calm neutral expression. This is the only readable face in
the image.

The person: [AGE, BUILD, HEIGHT IN CM, SHOULDERS, POSTURE]. [HAIR: LENGTH,
COLOUR, HOW IT SITS, PARTING, FINISH.] [BROWS, EYES, NOSE, FACE SHAPE, FACIAL
HAIR.] [ONE OR TWO DISTINGUISHING MARKS.] Real skin texture with visible pores,
natural asymmetry, fine lines, no retouching, no beauty filter, no digital
smoothing, no glossy or plastic skin.

Clothing: [EACH GARMENT - FABRIC, CUT, COLOUR, HOW IT HANGS, CONDITION]. [THEN
THE NEGATIONS: NOT TIGHT, NOT TECHNICAL, NO LOGOS, NO VISIBLE BRANDING.]
Footwear: [DESCRIBED IN FULL].

Panel 1 clearly shows [WHAT ONLY THE FRONT REVEALS]. Panel 3 shows [WHAT ONLY
THE BACK REVEALS]. All four panels matched exactly in framing, scale, anatomy,
clothing and lighting.

Look: plain solid #7f7f7f mid-grey seamless studio background, flat even
shadow-free studio lighting, matte fabrics with no sheen, 85mm lens, f/8,
photorealistic RAW detail. One person alone in frame - no props, no other
objects, no second figure, no text, no watermark, no panel numbers, no borders.

Stage 6: For products, never make the three-quarter your master

Product sheets follow the same one-image, four-panel logic, with different panels: the flattest informative view, the three-quarter, the surface the master hides, and one more that earns its place. The decision worth naming is which view goes first.

The three-quarter is the better marketing shot and the worse spec. It foreshortens, hides true proportion, and everything generated from it inherits that distortion permanently. Make the master the flattest orthographic view. For a shoe that is the direct side profile, which shows the curve of the sole and where every seam sits. There is a reason every footwear brand on earth uses the lateral profile as its primary image. The three-quarter goes second, where it adds depth once proportions are locked.

One more rule catches everyone. Any panel showing a surface the master hides has nothing to inherit, so it invents. A shoe underside hallucinates a tread pattern every time unless you describe it in full, negations included.

The product version of the same sheet. Note that the flat view goes first, not the three-quarter.

textFILL-IN RULE: Fill in the short list below. Leave blank anything you do not
know. Then write the finished prompt: take my answers and fold them into the
template underneath, so what comes back is one clean prompt with no brackets
left in it and no leftover list on top. Anything I left blank, work it out
from the rest of the prompt yourself. Do not ask me first. At the end, tell me
in two lines what you decided on my behalf so I can change anything you got
wrong.

OBJECT (what it is, in a word or two):
PURPOSE (what it is for - this is what makes the shape correct):
SILHOUETTE (proportions, stance, weight):
WHAT IT MUST NOT BE (three or four things):
MATERIALS (material, then construction, then hardware and closures):
COLOUR (positive colour names, and where each colour lives):
CONDITION (wear, creases, patina):
PANEL 1, THE MASTER VIEW (the flat straight-on view, never the three-quarter):
PANEL 2, THE THREE-QUARTER:
PANEL 3, THE HIDDEN SURFACE (the side the master view cannot show):
PANEL 4, THE FOURTH VIEW:
THE HIDDEN SURFACE IN FULL (a sole is flat, no tread, unless you say so):

--- THE PROMPT (leave this part alone) ---

A four-panel product sheet. This image contains four views in total: [PANEL 1],
[PANEL 2], [PANEL 3], [PANEL 4]. All four must be present and all four must show
the same single [OBJECT].

Panel 1 - [THE MASTER VIEW, DESCRIBED PRECISELY, WITH "NO PERSPECTIVE
DISTORTION" IF ORTHOGRAPHIC].
Panel 2 - [THE THREE-QUARTER, DESCRIBED PRECISELY].
Panel 3 - [THE HIDDEN SURFACE].
Panel 4 - [THE FOURTH VIEW].

The product: [WHAT IT IS AND WHAT IT IS FOR - THE PURPOSE LINE, WHICH IS WHAT
MAKES THE SHAPE CORRECT]. The silhouette is [PROPORTIONS, STANCE, WEIGHT]. NOT
[THREE OR FOUR THINGS IT MUST NOT BE].

Materials and colour: [MATERIAL, THEN CONSTRUCTION, THEN HARDWARE AND CLOSURES,
THEN COLOUR IN POSITIVE COLOUR NAMES, THEN CONDITION - WEAR, CREASES, PATINA].
[WHERE EACH COLOUR LIVES. WHICH PARTS ARE WHICH MATERIAL.] No logos, no
branding, no text anywhere on the [OBJECT].

[THE SURFACES THE MASTER DOES NOT SHOW, DESCRIBED IN FULL. A SOLE IS FLAT WITH
NO TREAD PATTERN, NO GROOVES AND NO TEXTURE UNLESS YOU SAY OTHERWISE.]

Look: plain solid #7f7f7f mid-grey seamless studio background, soft even
shadow-free lighting with a gentle top light to reveal surface changes, matte
with no gloss or reflections, 85mm lens, f/8, sharp commercial product
photography, photorealistic with fine material detail. One [OBJECT] alone -
no hanger, no stand, no hands, no other objects, no text, no borders.

Stage 7: Save it as an Element, or the tag will not work

A generated image is not yet an asset. An asset is a saved Element: a named, reusable reference living in the workspace. Saving it is what makes it addressable as an @tag in a scene prompt, and reusable across every scene after that. One Element per subject. The four-panel sheet is one Element, not four.

Name Elements by story role, never by description. Use @spy, @mark, @vendor. Do not use @woman or @person-1. The role tells the model how the character should carry themselves, which affects posture, gaze, and pace. The description tells it nothing it cannot already see in the image.

Stage 8: Write every asset line as take-this and ignore-that

In the scene prompt, every Element gets its own line, and every line says two things: what to take from the reference, and what to leave behind. The first half is obvious and everyone writes it. The second half is the one people forget, and forgetting it undoes the entire neutral-asset discipline at the very last step.

text@vendor - the vendor. Face, build and wardrobe exactly per ref. Do not
use the sheet's grey background or its flat lighting.

Without that second sentence, the neutral studio grey leaks straight into the finished shot. You kept the asset clean for the whole pipeline and then invited the studio into the bar. Every character and product line needs the background and lighting exclusion, without exception. Environment lines get the mirror version: take the space, materials, and lighting, ignore any people in the plate.

Stage 9: Write the look once and use it twice

The look block covers sharpness and format, then colour and grain, then how the light behaves, then the mood. Write it once. Then paste that same block, word for word, into every environment plate prompt and into the scene prompt. That shared sentence is the entire mechanism by which the plates and the finished film end up graded alike.

Two rules govern what goes in it. Derive the look from this subject, never from the last project. The default trap is warm amber tungsten with deep shadows, genuinely beautiful and genuinely wrong for about four scenes in five. And name colours positively. Avoid desaturated, monochromatic, and chiaroscuro, because those words return black and white footage more often than a grade.

The plate prompt. Whatever you write on the look line here is the line you paste into the scene prompt, word for word.

textFILL-IN RULE: Fill in the short list below. Leave blank anything you do not
know. Then write the finished prompt: take my answers and fold them into the
template underneath, so what comes back is one clean prompt with no brackets
left in it and no leftover list on top. Anything I left blank, work it out
from the rest of the prompt yourself. Do not ask me first. At the end, tell me
in two lines what you decided on my behalf so I can change anything you got
wrong.

LOCATION:
THE SPACE (architecture, scale, materials, surfaces, what is on the walls and
  floor, what the camera can see from here):
WHAT IS LIT INSIDE THE SPACE (lamps, windows, screens, signage, and where they
  sit):
THE LOOK (if you already have a LOOK block for this film, paste it here and it
  gets used word for word. If you leave it blank, write it from what this scene
  is about: sharpness and format, then colour and grain in positive colour
  names, then how the light behaves, then the mood. Do not reach for warm amber
  tungsten with deep shadows unless this subject really calls for it. Then tell
  me the LOOK you wrote, because it has to go into the scene prompt unchanged):

--- THE PROMPT (leave this part alone) ---

A single photograph of [THE LOCATION], empty of people. [THE SPACE:
ARCHITECTURE, SCALE, MATERIALS, SURFACES, WHAT IS ON THE WALLS AND FLOOR, WHAT
THE CAMERA CAN SEE FROM THIS POSITION.] [WHAT IS PRACTICAL AND LIT WITHIN THE
SPACE - LAMPS, WINDOWS, SCREENS, SIGNAGE - AND WHERE THEY SIT.]

[THE LOOK]

Completely empty - no people, no figures, no animals, no text, no watermark.

Stage 10: Choose Elements over a trained identity

There is a second path, worth naming so you rule it out deliberately rather than by accident. You can train an identity model from five to twenty photos of one person, and it produces a very strong likeness.

The deciding constraints are hard ones. A trained identity runs only on its own dedicated models, so it does not run on Seedance at all, and it permits exactly one identity per generation, which makes any two-hander impossible.

Stage 11: Read the failure modes before you hit them

Each of these has a specific cause. Match the symptom rather than rewriting the prompt from scratch.

A worked example: two people, one back room, one case

A vendor opens a case of odd goods on a counter, a customer deadpans a line, then reveals a stranger case of his own. The scene needs two people, one prop, one location, and it runs about fifteen seconds. Here is what gets built before a word of the scene is written.

Two character sheets, four panels each on flat mid-grey, heads turned down and away on the front panels. One product sheet for the case, mastered on its flattest face, with the interior described in full because no other panel shows it. One plate of the back room, empty of people, generated in the film's actual look with the practical lamps where they really sit. That comes to four generations and four Elements, named @vendor, @customer, @case and @counter-room.

And the scene itself, once every asset above is saved as an Element.

textFILL-IN RULE: Fill in the short list below. Leave blank anything you do not
know. Then write the finished prompt: take my answers and fold them into the
template underneath, so what comes back is one clean prompt with no brackets
left in it and no leftover list on top. Anything I left blank, work it out
from the rest of the prompt yourself. Do not ask me first. At the end, tell me
in two lines what you decided on my behalf so I can change anything you got
wrong.

WHAT HAPPENS (one sentence, start to finish, what people do and what they want):
TOTAL DURATION IN SECONDS:
LOCATION (and its @tag):
WHO IS IN IT (one line each: their @tag, and who they are in the story):
WHAT THEY WEAR AND CARRY (anything that has to hold from start to finish, and
  anything that changes and when):
PROPS (@tag, and what the object is):
DIALOGUE (who speaks and their exact words, or none):
VOICE REFERENCE (@tag, and which character it belongs to):
THE LOOK (paste the LOOK from your environment plate word for word):
CUTTING RHYTHM (how fast the scene should cut and how it should feel):
SOUND (music and where it drops out, named effects, ambience underneath,
  subtitles on or off):
ANYTHING YOU DO NOT WANT IN FRAME:

--- THE PROMPT (leave this part alone) ---

[PREMISE]
[WHAT KIND OF SCENE THIS IS AND WHAT HAPPENS START TO FINISH, IN ONE SENTENCE.
MEANING, NOT CAMERA - WHAT PEOPLE DO AND WHAT THEY WANT.] [TOTAL DURATION IN
SECONDS.]
[ASSETS]
@[CHARACTER-ROLE] - [WHO THEY ARE IN THE STORY]. Face, build and wardrobe
  exactly per ref. Do not use the sheet's grey background or its flat lighting.
@[SECOND-CHARACTER-ROLE] - [WHO THEY ARE IN THE STORY]. Face, build and wardrobe
  exactly per ref, including [ANY ITEM THAT MUST CARRY OVER]. Do not use the
  sheet's background or lighting.
@[PROP-NAME] - [WHAT THE OBJECT IS]. Structure, proportions, materials and
  colour exactly per ref.
@[LOCATION-NAME] - the location. Space, materials and lighting only. Do not use
  any people in the image.
@[AUDIO-NAME] - voice reference for [WHICH CHARACTER]. Match this voice exactly
  for their dialogue. Do not use any other audio from it.
[LOCKS]
[CHARACTER-ROLE]: [WARDROBE THAT HOLDS THROUGHOUT], [HOW THEY CARRY THEMSELVES].
[SECOND-CHARACTER-ROLE]: [WARDROBE THAT HOLDS THROUGHOUT], [ANYTHING THAT
CHANGES AND THE EXACT BEAT IT CHANGES AT]. One location and one lighting setup,
start to finish. [WHO STAYS FRAME LEFT AND WHO STAYS FRAME RIGHT.]
[BEATS]
[BEAT 1 | 0-2s | NAME OF THE BEAT] [PHYSICAL ACTION AND STAGING ONLY. WHAT IS
VISIBLY TRUE WHEN IT FINISHES - WHERE EACH PERSON IS, WHO IS HOLDING WHAT.] NO
CUT.
[BEAT 2 | 2-4s | NAME OF THE BEAT] [PHYSICAL ACTION. END STATE.] CUT.
[BEAT 3 | 4-6s | NAME OF THE BEAT] [PHYSICAL ACTION. END STATE.] CUT.
[KEEP GOING UNTIL THE RUNTIME IS FULL. A SPOKEN LINE RUNS 2-3 SECONDS, SO COUNT
THE DIALOGUE FIRST. DIALOGUE IN ORDINARY DOUBLE QUOTES.]
[LOOK]
[SHARPNESS AND FORMAT.] [COLOUR AND GRAIN, IN POSITIVE COLOUR NAMES ONLY - NAME
THE COLOURS YOU WANT, NEVER "DESATURATED", "MONOCHROMATIC" OR "CHIAROSCURO".]
[HOW THE LIGHT BEHAVES, INCLUDING ANY FLICKER OR MOVEMENT WRITTEN OUT
EXPLICITLY.] [THE MOOD.]
[DIRECTION]
[CUTTING RHYTHM.] [CAMERA MOVES YOU DO NOT WANT.] [ONE LINE PER CHARACTER ON HOW
THE PERFORMANCE SHOULD READ.] [ONE CAMERA ANGLE ONLY IF ONE SHOT HAS TO BE A
REVEAL - OTHERWISE LET THE MODEL CHOOSE COVERAGE.]
[SOUND]
[MUSIC AND THE BEAT IT DROPS OUT ON.] [NAMED EFFECTS AND THE BEATS THEY HAPPEN
ON.] [AMBIENCE UNDERNEATH.] [SUBTITLES ON OR OFF.]
[EXCLUDE]
No brand logos, no extra people, no on-screen text. [ANYTHING ELSE YOU DO NOT
WANT IN FRAME.]

Now the scene prompt is mostly pointing. Each Element gets its take-and-ignore line. Screen geography gets locked once, so the vendor stays frame left and reverses stop flipping sides. The beats get timed against the dialogue at roughly two to three seconds per spoken line. The look block is pasted in identically to the one already used on the plate. The prompt is long, but almost none of it is description, because the description was done days ago.

Stage 12: The path from reading this to having a scene

The shortest honest sequence from finishing this page to a finished shot, in one sitting.

  1. List every person, object, and location the scene needs. That list is your build order.
  2. Write the look block first, before any asset, because the plates need it verbatim.
  3. Generate each character sheet at 2K: four panels, the neutral grey line, count stated first, one readable face.
  4. Generate each product sheet the same way, mastered on the flattest orthographic view, hidden surfaces described.
  5. Generate one plate per location, empty of people, with the look block pasted in unedited.
  6. Approve each sheet at full size. Check the panel count, the face count, and that no colour crept into the background.
  7. Save each approved image as an Element, named by story role.
  8. Write the scene prompt: premise, one asset line per Element with both halves, then the locks, the beats, and the look block.
  9. Render a cheap low-resolution test. Watch for grey leak, face drift, and side flips.
  10. Fix what the test showed, then render at full resolution.

Common questions

Do I really need a character sheet if the person only appears in one shot?

No. A single shot with a single appearance does not need one, and building a sheet for it is wasted time. The sheet earns its cost the moment the same person has to appear twice, because that is when drift becomes visible.

Can I use a real photo of someone instead of generating a character?

Yes, as the input to the sheet rather than as the asset itself. Feed the photo in and generate the four-panel sheet from it, so you still end up with the neutral grey field, the four views, and the single readable face. The photo alone carries its own lighting, which is exactly what the neutral-asset rule exists to strip out.

Why not just describe the character in every scene prompt instead?

Because description is not deterministic. The same paragraph produces a slightly different person each time, and the differences compound across a cut. A saved Element is a fixed thing to point at, which is a different mechanism entirely.

What if I want the character lit warmly in the final scene?

You will get it. The warmth comes from the environment plate and the look block, not from the sheet. Keeping the sheet neutral is what lets you light the same character warmly in one scene and coldly in the next without rebuilding anything.

How many Elements is too many for one scene?

There is no hard ceiling, but every Element in a prompt is another thing the model has to reconcile. If a scene needs more than about five, check whether some of the props really need to be assets or whether they can just be described in the action.

Does this work the same way for a single continuous take?

The asset stage is identical. Only the scene prompt changes: drop the numbered beats, state that it is one continuous take with no cuts, and write the action as a single flowing paragraph. That version tends to be better for handheld and phone footage, where the camera is meant to read as a person.

The prompts are yours. The system is in the club.

Everything on this page you can run by hand, one prompt at a time. Inside the club there are two skills that do it for you: one builds the sheets and plates and saves them as tagged Elements, the other takes an idea and returns a finished scene.

Join the Club — $9/mo

Read this online at claudecodeclub.ai