Seedance 2.5 Full Step-By-Step Guide + Where To Find It

@maxxmalist
MAX@maxxmalist
213 views Aug 06, 2026 ~11 min read
Advertisement

seedance 2.5 is coming soon on Higgsfield

i went through every doc, every spec and every workflow, here's the full breakdown so you're ready before it drops 🧵

Media image

what even is seedance 2.5

it's the direct upgrade to seedance 2.0, the model that already changed how AI ads get made

2.0 gave you the reference system. 2.5 takes the same idea and scales it to a level nobody else is even attempting:

30-second videos in ONE generation. up to 50 reference files at once. region-level editing so you fix one detail instead of re-rolling the whole clip

you're not generating clips anymore. you're producing entire ads in a single pass

how you'll access seedance 2.5

remember the 2.0 launch? chinese phone number workarounds, sms verification sites, region locks. none of that this time

seedance 2.5 is coming soon to Higgsfield:

  • go to higgsfield ai
  • create an account now
  • when it drops, you select Seedance 2.5 as your model and you're in
  • no region locks, no workarounds. the model everyone gatekept last time, one click away. everything in this guide is the workflow you'll be running inside Higgsfield from day one

    the specs (2.0 vs 2.5)

    this is not a small update:

  • images: 9 -> 30 (up to 4K resolution each)
  • videos: 3 -> 10 (30s max total upload)
  • audio: 3 -> 10 clips (30s max total, and audio-only reference is now supported)
  • output duration: 15s -> up to 30 seconds in a single continuous generation
  • native sound, lip-sync, and music generation built in
  • 10+ languages: english, spanish, portuguese, arabic, japanese, korean, thai, vietnamese, indonesian, malay + chinese
  • up to 50 total reference slots. you can hand it complete character sheets, scenes, live footage, storyboards AND audio all at once, and it actually remembers the structure and logic behind every file. your entire cast, props and style refs travel together in one generation

    getting started

    once seedance 2.5 lands on Higgsfield, the flow looks like this:

    select Seedance 2.5 as your model
    upload your assets, they get labeled automatically (@Image1, @Video1, @Audio1 etc)
    assign roles in your prompt: "@Image1 is the main character, reference @Video1 for camera movement, use @Audio1 for pacing"
    pick aspect ratio (9:16 tiktok, 16:9 youtube) and duration (up to 30s)
    generate

    the @ system from 2.0 is still the core mechanic. every file gets a job. you're directing, not prompting

    the prompt formula that actually works

    2.5 responds insanely well to structure. this is the exact format:

    [creatives description] + [one-sentence summary] + [timestamped plot] + [global rules]
  • creatives description = what each @ file is for (character / voice / action / scene). never leave a file unassigned
  • one-sentence summary = subject + location + event + style + camera
  • timestamped plot = cut your 30 seconds into slices: "0-8s: …", "9-16s: …", each with picture + camera + action + dialogue + sound
  • global rules = what must stay consistent + what's banned ("no subtitles, no bgm, no fast cuts")
  • the timestamp storyboard is the single biggest unlock. you're literally writing a shooting script and the model follows it second by second

    real example: 30-second UGC ad in one generation

    here's exactly what this looks like in practice. two reference images in, a full ad out in a single pass

    reference images:

    Media image

    the exact prompt:

    UGC selfie-style skincare video, shot on a smartphone front camera held at arm's length by the subject herself. Vertical 9:16 framing, natural handheld phone-camera imperfections: slight shake, casual autofocus hunting, everyday indoor bathroom lighting, no color grading, no cinematic polish. This must NOT look like a commercial, it must look like a real person filming a casual 30-second selfie video in her own bathroom, one continuous handheld take, no professional camera movement, no dolly, no gimbal smoothing, no edited cuts between shots. Actions should be slow, deliberate, and unhurried, this is a calm real person, not a fast edit. [Multi-Modal Reference Layer] Based on uploaded @image1, strictly keep the young woman's face, hairstyle, skin tone, features, pose, outfit, and bathroom identical to the first frame, this is the exact same person, room, and starting moment as the reference image, nothing is in her free hand yet. Based on uploaded @image2, strictly keep the product identical: a flat lavender-and-cream single-use sheet mask sachet with a thin gold divider line, serif logo reading "Lumea Glow Ritual", clean skincare-brand design, no other logos. [Character Styling] [Age] Woman in her early-to-mid 20s, natural fresh-faced look, no heavy makeup. [Skin] Warm light-tan skin with natural glow, light freckles across the nose and cheeks, realistic pores, no retouching. [Facial Features] Full lips, softly defined brows, relaxed calm expression. [Eyes] Warm brown eyes, steady relaxed gaze, direct eye contact with the lens as if talking to a friend, glancing down when handling the product. [Hair] Long dark brown hair, slightly tousled, loose strands around the face. [Clothing] Olive-green ribbed tank top, light grey sweatpants, thin gold necklace, exactly as in @image1. [Framing] Extreme selfie framing, arm extended holding the phone toward herself, standing in the bathroom exactly as in @image1. [Global Setting] A bright modern bathroom matching @image1 exactly: warm cream walls, white ladder-style towel radiator behind her, open sliding door on the left leading to a darker hallway, white bathtub edge in the lower right, recessed ceiling spotlight. Soft warm indoor lighting, no studio lighting. [Camera Language] Handheld selfie phone framing for the full 30 seconds, held slightly below eye level tilted up, natural micro-shake throughout, no smooth gimbal movement, no locked tripod shots, occasional soft refocus blur when an object gets close to the lens, single unbroken take with no hard cuts. [Continuity] Once an object appears in frame (towel, sachet, scissors, sheet mask), it stays consistent, nothing appears or disappears without a visible hand action bringing it in or setting it aside off-frame. After the mask is applied, it stays on her face for the rest of the video, not shifting between moments. [Speech Direction] She talks less overall, with real pauses, not narrating constantly. Natural hesitation: filler words like "uh" and "um", brief mid-sentence pauses, trailing off instead of finishing every sentence cleanly, occasional small self-correction, quiet moments where she's just doing the action without talking. Soft casual tone, slightly tired end-of-day energy, like someone thinking out loud, not reading a script. [Timestamp Storyboard] 0s-3s: She stands exactly as in @image1, looks into the lens, a small beat, then says casually: "Okay so, um... night routine. Kind of." Small half-smile. 4s-7s: Quiet, no dialogue. She reaches off-frame, brings a soft white towel into view, gently pats her face dry, slow and unhurried, then sets the towel off-frame. 8s-12s: She reaches off-frame and brings the sachet from @image2 into view, holds it up near the lens so it goes softly out of focus for a second, says: "This is the one everyone's been, like... posting about. So." Small shrug. "We'll see." 13s-17s: She reaches off-frame for small scissors, snips straight across the top edge of the sachet, one slow clean cut, sets the scissors aside off-frame. Just the crisp snip sound, says quietly: "I always rip these wrong, so... scissors." 18s-24s: She pulls out the folded sheet mask, glistening with essence, unfolds it, says: "Okay it's... very wet. Good sign, I guess." Then slowly lays it onto her face starting from the forehead, smoothing it over cheeks and chin, adjusting the eye holes, phone dips slightly, a soft laugh: "Hold on... okay." 25s-27s: Quiet, no dialogue. Mask fully on, she smooths the edges by her jaw, looks at herself through the lens like a mirror. 28s-30s: She looks into the lens with the mask on, a small relaxed smile, a pause, then: "Fifteen minutes and, uh... we'll see if it's worth the hype." Beat. Slow exhale. [Global Supplement] Audio: only her own spoken voice with natural breath sounds and pauses, natural room tone (towel patting, crisp scissor snip, wet rustle of the sheet mask, light foil crinkle), no background music, no subtitles, no voiceover from someone else

    the result:

    0:30

    two images, one prompt, one generation. consistent face, consistent product, synced dialogue with native lip movement, 30 seconds of usable ad. this workflow used to take me 4-6 separate generations plus editing. and this is exactly what you'll be doing with seedance 2.5

    what you can actually do

    30-second one-take narratives - complete emotional arcs with zero post-production stitching. write a 5-stage storyboard (setup -> tension -> turn -> release -> close) and it plays the whole thing in one continuous shot. i've seen a 29-second close-up of an actress going from a question to tears to a smile, all directed by timestamps

    physics that hold up - sand, dust and debris react to every impact. water splashes obey real fluid dynamics, fabric drapes and flows with actual weight, tires kick up spray that behaves like spray. this is the difference between "AI video" and footage

    camera moves that feel operated, not simulated - crash zooms and whip pans land right on the action. dolly-ins hit their mark, orbits stay locked on the subject, handheld breathing feels like a human holding the rig. you write the camera language, it executes like an operator who read the shot list

    characters don't drift - face identity locked across all 30 seconds, and blocking works better than any other generation of the model: multi-character scenes hold their spatial positions shot to shot, nobody teleports, nobody morphs mid-clip

    lighting stays motivated - light sources make sense from every angle. candlelight flickers on the correct side of the face, window light stays consistent when the camera cuts, rim lighting follows the subject through the move. no more scenes where the sun jumps around between shots

    50-reference character + scene lock - upload a full character sheet, product shots, the location, a motion reference and a voice sample in one generation. best practice from the internal docs:

  • 1-5 subjects = most stable results
  • 5-10 second subject clips work better than long ones
  • separate images beat one combined multi-view image
  • region-level editing (the biggest one) - draw a box, arrow or brush stroke directly on the video and tell it what to change: "inside the red box, replace the boy's jeans with black suit pants, from 0s to 8s" change one element, keep everything else identical. no more re-rolling a 90% perfect clip because of one wrong detail

    clay renderer / white model workflow - feed it a rough 3D blockout video and it renders your characters onto the exact motion and camera path. "replace the white model in @Video1 with the character in @Image1, keep all camera movement" = frame-accurate control that pro studios pay previz teams for

    green screen editing - upload green screen footage, swap backgrounds, composite like a post house

    video extension with real transitions - extend any clip forward or backward and specify the exact transition: match cut, whip pan, mask transition, ink-wash bleed, dynamic relay (character jumps out of frame A, lands in frame B wearing different clothes). you name the film-school technique, it executes it

    10-language lip-sync - one winning creative, ten localized versions with native mouth movement. no reshoots, no dubbing budget

    how you'll use it:

  • ads & ecommerce = one 30s VSL in a single generation, winning creatives replicated with your product via reference stacking, region edits to test hook variations without regenerating
  • content localization = same ad, 10 markets, native lip-sync in each
  • short dramas & storytelling = timestamped multi-shot scripts with consistent characters across the full 30 seconds
  • previz & client work = white model blockouts turned into finished renders, exact camera and staging control without third-party 3D software
  • ai influencers = 30-second talking clips with locked identity across every video
  • pro tips

  • write the narrative arc FIRST, then shot details. opening -> progression -> turn -> resolution beats stacking visual adjectives
  • layer your references: characters, product, scene, rhythm, audio as separate assets instead of one messy dump
  • use the real-person character formula for humans: age/ethnicity + skin texture ("retain real micro pores and skin texture") + 3-4 facial details + what the eyes convey + hair state + fabric texture + body/aura. this one formula erases the AI face
  • negative prompts are load-bearing at 30s: "no subtitles, no bgm, no fast cuts, no exaggerated movements" written in global rules AND repeated at the end
  • for edits, always include timing: "only between 5s and 8s" prevents the change bleeding across the whole clip
  • keep edit-target videos under 20 seconds for max stability
  • @ your references multiple times throughout the prompt. repetition = precision
  • seedance 2.5 turns directing into full production, script, staging, shooting, editing and localization inside one model. and it's all landing on Higgsfield soon

    the people who learn this workflow before it drops will be shipping while everyone else is still reading the docs

    bookmark this. you'll need it.

    and for more free value join here: https://t.me/maxxmalist

    higgsfield is sponsoring this article.

    Actions
    What You Can Do
    • Export as PDF or Markdown
    • Batch Export to Notion
    • Bookmark & Highlight
    • LinkedIn & Instagram Carousel Maker
    Create Free Account

    Includes 7-day Premium trial

    Advertisement