We made the first 110-minute AI feature film with a real cast, The...

Starring @N3onOnYT, @stylebender, @RampageJackson, and @MKIATPIS
It's 100% open-sourced on Higgsfield: all prompts and assets are public now.
Made with Seedance on Higgsfield.
Watch full film below 👇
You have access to $2,000,000 worth of AI filmmaking knowledge.
higgsfield.ai/@higgsfield.st…
We digitised their likeness: the character sheets were built from contract photography, so the face on the sheet is the actor's face - not a "similar type". After that, only the sheet goes into a shot.
Likeness and voice rights were closed before the first generation.
Take the head off the full-body figures. On the wide panels the face is small and soft - and that's exactly the face the model will copy into a wide shot, badly. Remove it, and there is only one place left to take a face from: the close portrait.
Make two close-ups - with a smile and without. Otherwise the model invents the teeth, and the smile arrives as somebody else's mouth.
Every asset passed a stress test before it was locked: ten generations in different poses and light, recognisable in ten out of ten. If the test fails, the problem is your description, not the model.
Splitting is cheaper than arguing.
The block is pasted into the audio field as is, every time the hero speaks, and it never changes. Not even a synonym - changing the wording widens what the model samples from, and the voice drifts.
The accent is written out phonetically inside the line: th going to f and v, dropped h, glottal t, -ing to -in'. One character, one accent, and it never drifts in any shot.
Both bosses live in red and gold, but the old one is patina - dried oxblood, tarnished gold, one aged lamp - and the new one is polish: saturated crimson, mirror-bright brass, identical lamps in symmetry.
Never tarnish on the new money, never mirror polish on the old. Written that way, two rich interiors stop looking like one set.
Not frontal: a frontal picture of a room is flat wallpaper — the model cannot read volume from it, and past the frame edges it invents new surroundings every time. Leave an anchor in every location and tie the staging to it: "the hero at the lamp, facing the door" works, "the hero in the room" is a lottery.
The trick we found late: generate a video of the empty location, the camera slowly walking through - the model draws the other sides of the room to match your plate. A full location kit out of one single image.
Every shot is born from references plus text: the assets carry the picture, and the wording carries everything else - including the geometry.
It is harder - and it is what makes the shots cut together.
The speech count: the take contains exactly three words - "Pull it, Oli." - written as its own lock. Without it the model adds a mumble or a line in another language; it does not like silence.
The height ruler: not "taller" - "Horace's eye-line at the level of Cal's mouth". The off-screen event: the door Oli kicks is written as a list of what must not appear - not a limb, not a shadow, not a reflection.
Different rhythms: two people react to one event, never in sync. A synchronised reaction reads as animation.
So the music never comes from the model. The track is recorded first, cut into 12-second blocks on the vocal's breaths, and each block goes in as a file: THE TRACK THEY ARE PERFORMING - the audio of this file IS the live performance.
Our earlier version explained the workflow - "audio guide, for sync only". Every word of it was true, and it produced nothing. The working version says the simple thing: this is the song, and he is singing it.
And mouth ownership has to be assigned, or everyone in frame starts mouthing the verse - a chorus of ventriloquists.
The fight inside the Cadillac would not come together: two bodies in constant contact, a struggle over a gun - limbs fused, the weapon vanished and reappeared in the other hand. So we called in stunt performers, choreographed the fight, and filmed the takes in a real car on an iPhone.
That footage went into the generation as motion reference - the model laid our characters, our interior and our light over it. One day of phone footage instead of a week of iterations. The only place we step outside the text-to-video-only rule.
Ours has 137 entries, and it shows which shots fought back: v15, v10, v9. Without the log you cannot repeat a good shot, and you cannot tell whether you already tried this fix.
The ten-to-fifteen rule holds: not there after that many iterations - the problem is not the wording. Simplify the shot: split it in two, remove an action, change the angle.
1. Assets first. Nothing generates until every face, place and prop is locked.
2. Describe everything, every time. The model has no memory.
3. Change one thing at a time.
4. Give the model less freedom - a corner, not a room; a map, not guesswork.
5. Shot won't land? Simplify the shot, not the words.
Every rule exists because a shot failed without it.
Everything that built the film is public - take it and make yours.
higgsfield.ai/@higgsfield.st…
