114:52

We made the first 110-minute AI feature film with a real cast, The...

@higgsfield_ai
Higgsfield AI 🧩@higgsfield_ai
42 views Aug 11, 2026 ~6 min read
We made the first 110-minute AI feature film with a real cast, The Cully Hill Boys, on Higgsfield for $2,000,000.

Starring @N3onOnYT, @stylebender, @RampageJackson, and @MKIATPIS

It's 100% open-sourced on Higgsfield: all prompts and assets are public now.

Made with Seedance on Higgsfield.

Watch full film below 👇
The short version is below - all the assets, prompts, and a much more detailed brief are on the project page.

You have access to $2,000,000 worth of AI filmmaking knowledge.

higgsfield.ai/@higgsfield.st…
First thing worth doing: grab the CINEDANCE skill and open the canvas, this is where we keep all the sheets, locations, etc.
The film has signed actors in it. Oli is Israel Adesanya, the UFC fighter. Tobin is Quinton "Rampage" Jackson. Horace is N3on, one of the loudest streamers around.

We digitised their likeness: the character sheets were built from contract photography, so the face on the sheet is the actor's face - not a "similar type". After that, only the sheet goes into a shot.

Likeness and voice rights were closed before the first generation.
A character sheet is three panels: a full body from the front, a full body from the back, and a large close portrait in 3/4 view.

Take the head off the full-body figures. On the wide panels the face is small and soft - and that's exactly the face the model will copy into a wide shot, badly. Remove it, and there is only one place left to take a face from: the close portrait.

Make two close-ups - with a smile and without. Otherwise the model invents the teeth, and the smile arrives as somebody else's mouth.
A hero has as many assets as states he goes through. Cal at the start - clean jacket, orange rucksack. Cal after the Thames - wet. Cal in the third act - split brow, blood on a white tee. Not one asset with a note: three different assets. Mix them in one text, and the blood fades in and out, the jacket comes back.

Every asset passed a stress test before it was locked: ten generations in different poses and light, recognisable in ten out of ten. If the test fails, the problem is your description, not the model.

Splitting is cheaper than arguing.
The voice is not an asset - it's a set of precisely written conditions: register, tempo, accent, manner.

The block is pasted into the audio field as is, every time the hero speaks, and it never changes. Not even a synonym - changing the wording widens what the model samples from, and the voice drifts.

The accent is written out phonetically inside the line: th going to f and v, dropped h, glottal t, -ing to -in'. One character, one accent, and it never drifts in any shot.
The year is not decoration, it is a rule: nothing in frame is newer than 2011. No smartphones, no glowing screens in a crowd, no modern cars at the kerb. The model pulls the picture toward today by default - give it room, and an extra will be holding an iPhone.

Both bosses live in red and gold, but the old one is patina - dried oxblood, tarnished gold, one aged lamp - and the new one is polish: saturated crimson, mirror-bright brass, identical lamps in symmetry.

Never tarnish on the new money, never mirror polish on the old. Written that way, two rich interiors stop looking like one set.
Generate locations for your future camera angles.

Not frontal: a frontal picture of a room is flat wallpaper — the model cannot read volume from it, and past the frame edges it invents new surroundings every time. Leave an anchor in every location and tie the staging to it: "the hero at the lamp, facing the door" works, "the hero in the room" is a lottery.

The trick we found late: generate a video of the empty location, the camera slowly walking through - the model draws the other sides of the room to match your plate. A full location kit out of one single image.
One condition the whole film stands on: text-to-video only. No starting frame, no image-to-video.

Every shot is born from references plus text: the assets carry the picture, and the wording carries everything else - including the geometry.

It is harder - and it is what makes the shots cut together.
Four devices from one corridor prompt, worth stealing.

The speech count: the take contains exactly three words - "Pull it, Oli." - written as its own lock. Without it the model adds a mumble or a line in another language; it does not like silence.

The height ruler: not "taller" - "Horace's eye-line at the level of Cal's mouth". The off-screen event: the door Oli kicks is written as a list of what must not appear - not a limb, not a shadow, not a reflection.

Different rhythms: two people react to one event, never in sync. A synchronised reaction reads as animation.
A video model will not perform your song. Ask it to "rap" and you get a mouth moving to nothing.

So the music never comes from the model. The track is recorded first, cut into 12-second blocks on the vocal's breaths, and each block goes in as a file: THE TRACK THEY ARE PERFORMING - the audio of this file IS the live performance.

Our earlier version explained the workflow - "audio guide, for sync only". Every word of it was true, and it produced nothing. The working version says the simple thing: this is the song, and he is singing it.

And mouth ownership has to be assigned, or everyone in frame starts mouthing the verse - a chorus of ventriloquists.
When a stunt will not generate - shoot it yourself.

The fight inside the Cadillac would not come together: two bodies in constant contact, a struggle over a gun - limbs fused, the weapon vanished and reappeared in the other hand. So we called in stunt performers, choreographed the fight, and filmed the takes in a real car on an iPhone.

That footage went into the generation as motion reference - the model laid our characters, our interior and our light over it. One day of phone footage instead of a week of iterations. The only place we step outside the text-to-video-only rule.
We generated in batches, scene by scene. Every iteration was surgical: one line changes, the rest stays word for word. And everything goes into the log - version, what changed, verdict.

Ours has 137 entries, and it shows which shots fought back: v15, v10, v9. Without the log you cannot repeat a good shot, and you cannot tell whether you already tried this fix.

The ten-to-fifteen rule holds: not there after that many iterations - the problem is not the wording. Simplify the shot: split it in two, remove an action, change the angle.
The whole film runs on five rules:

1. Assets first. Nothing generates until every face, place and prop is locked.

2. Describe everything, every time. The model has no memory.

3. Change one thing at a time.

4. Give the model less freedom - a corner, not a room; a map, not guesswork.

5. Shot won't land? Simplify the shot, not the words.

Every rule exists because a shot failed without it.
This is a $2,000,000 production, and it is now open source.

Everything that built the film is public - take it and make yours.

higgsfield.ai/@higgsfield.st…
Actions
Advertisement