video model · by ByteDance
Seedance 2.5 AI Video Generator
Generate clips up to 30 seconds long with Seedance 2.5, ByteDance's newest video and audio model. Plan several connected shots in one take, guide it with up to 30 reference images, and get synchronized sound in the same pass. Available in Fliki's AI Playground. Compare it side by side with all our AI video models before you render.
Generated with Seedance 2.5
Seedance 2.5 clips generated inside Fliki. No edits, no post.
Prompt
A quiet short-film moment in a roadside diner after midnight, told in four shots with one exchange of dialogue, built to show off a single continuous performance carried across cuts. THE SUBJECT: Loretta, a Black American waitress in her late fifties with short silver-streaked natural hair, reading glasses pushed up on her head, a pale mint diner uniform dress with a white collar, a name tag with no readable text, and a small gold cross necklace. She moves with the ease of someone who has worked this counter for thirty years. Across the counter sits Danny, a white trucker in his early thirties with a sunburned neck, a sandy three-day beard, a faded red flannel shirt over a gray T-shirt, and a sweat-stained tan cap he takes off and holds in both hands. Props that stay consistent: a chipped white ceramic mug, a glass coffee pot two thirds full, and a slice of cherry pie on a plain white plate with a single fork. SETTING: A small 1970s roadside diner off an interstate in rural Nebraska at 1 a.m. Red vinyl stools, a long speckled Formica counter, chrome napkin dispensers, a pie case with a glowing lid, and a wide front window showing an empty highway, one blinking yellow traffic light, and a parked semi truck with its running lights on. Rain has just stopped; the window glass carries fine droplets. LIGHTING: Warm fluorescent tubes over the counter, slightly green in the shadows, mixed with the cool blue spill of the night outside the window and the amber pulse of the traffic light. Soft practical glow from the pie case lights the lower half of faces. Gentle contrast, deep but readable shadows, a subtle film grain. CAMERA: Shot on a 35mm cinema camera with vintage spherical lenses, shallow depth of field, slow deliberate moves. Mostly locked-off or slow push-ins at counter height. Natural color grade leaning warm, with the blacks lifted slightly like 35mm print film. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Medium shot from behind the counter at shoulder height. Loretta tops off the chipped mug in front of Danny; coffee streams in a steady arc and steam curls up. Danny sits with his cap in his hands, eyes on the pie. 2.5s to 5.0s: Close-up on Danny, slightly low angle, window and blinking light soft behind him. He looks up at her and speaks, voice tired and a little embarrassed: "I cannot pay for the pie." His lips sync to every word; he gives a small apologetic shrug. 5.0s to 7.5s: Close-up on Loretta, reverse angle matching his eyeline. A slow smile starts at the corners of her eyes. She sets the coffee pot down on the counter with a soft clink and answers warmly: "Pie is on me, honey." Her lips sync precisely. 7.5s to 10.0s: Wide shot from the far end of the counter, the whole diner visible. Danny lets out a breath, sets his cap on the stool beside him, and picks up the fork. Loretta walks away along the counter toward the kitchen pass, wiping her hands on a towel. The semi truck idles outside the window. SOUND: Continuous diner room tone: the hum of fluorescent tubes, a refrigerator compressor in the pie case, a distant truck passing on wet asphalt outside. Coffee pouring into the mug, the ceramic clink of the pot on Formica, the scrape of a fork on a plate at the end. A small radio in the kitchen plays an indistinct old country song very low, no clear lyrics. Dialogue is dry and close, with the room tone steady underneath and no music swell. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A tense sports moment at a small-town rodeo, four shots that move from calm preparation to explosive release, built to show how the model holds a rider and horse consistent through fast action. THE SUBJECT: Maya, a Mexican American barrel racer in her early twenties with a long dark braid down her back, a cream felt cowboy hat, a turquoise long-sleeve western shirt with pearl snaps, dark blue jeans, brown leather chaps, and worn brown boots. Her horse is a sorrel quarter horse mare named Pepper with a white blaze down her face, a black leather bridle with a silver concho, and a turquoise saddle pad that matches the shirt. The same horse and tack appear in every shot. SETTING: An outdoor county fair rodeo arena in west Texas at golden hour. A fenced dirt arena with three bright blue barrels, metal bleachers with a small crowd in the far background, the alley gate where riders wait, dust hanging in the air, and flat ranch land with a single windmill on the horizon. LIGHTING: Low late-afternoon sun from behind the rider, creating a warm rim light on the hat brim, the braid, and the horse mane. Dust in the air catches the light and glows. Faces are filled by soft bounce from the pale dirt. Rich warm tones, golden skin, deep blue sky. CAMERA: Sports documentary style. The first shots are handheld with a long lens and shallow depth of field; the release is a fast tracking shot. Slight natural camera shake, no stabilization wobble artifacts. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Tight close-up in the alley. Maya leans forward and presses her forehead to the mare's neck, eyes closed. She whispers softly, lips moving in sync: "Easy, girl. Just like home." Pepper flicks an ear back toward her voice. 2.5s to 4.5s: Extreme close-up of Maya's gloved hand shortening the leather reins, knuckles tightening, then her boot heel settling into the stirrup. Pepper paws the dirt once and a small puff of dust rises. 4.5s to 7.5s: Low wide tracking shot beside the arena fence. Pepper bursts out of the alley at full gallop toward the first blue barrel, dirt kicked high behind her hooves, Maya low over the saddle horn, braid flying. 7.5s to 10.0s: Medium shot as Pepper turns tight around the first barrel, leaning hard, Maya's inside hand on the horn, the barrel rocking but not falling. Dirt sprays toward the camera. The crowd in the far bleachers rises. SOUND: Alley shot: close, intimate sound of the horse breathing and snorting, leather creaking, the jingle of the bit, Maya's whisper. Then hoofbeats thunder in rhythm as the mare launches, dirt spraying, the rattle of the barrel as she turns, and a rising cheer from a small crowd with a few whistles. A distant arena announcer's voice is muffled and unintelligible. No music. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A premium product film for a stainless steel dive watch, four shots that alternate macro detail with a person wearing it underwater and on a boat, built for a vertical ad with one short spoken line. THE SUBJECT: The product: a stainless steel dive watch with a 42mm brushed case, a deep green sunburst dial, a rotating ceramic bezel in matching green with luminous markers, white luminous hands, and a brushed steel bracelet. No brand name or text is visible on the dial. The wearer is Kai, a Native Hawaiian free diver in his thirties with short black hair, a trimmed beard, a sun-darkened complexion, a simple black wetsuit top unzipped to the chest, and the watch on his left wrist. SETTING: Clear turquoise water off the coast of Maui over a white sand bottom with scattered lava rock, then the teak deck of a small white dive boat, wet from spray, with coiled rope and a folded mask resting on the rail. LIGHTING: Underwater: sunlight cuts down in moving caustic patterns across the sand and across the watch, blue-green color, soft light shafts. On deck: bright late-morning sun with crisp specular highlights sliding across the brushed steel and the domed crystal, deep blue sky behind. CAMERA: High-end commercial look. Macro probe lens for the product, slow motorized sliders, and a stabilized underwater housing. Crisp focus on the dial, gentle focus falloff everywhere else, luxurious slow motion feel without actual frame skipping. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Macro shot on deck. The watch rests on wet teak; a single water droplet rolls across the domed crystal while the camera glides left, a highlight sweeping across the green sunburst dial. 2.5s to 5.0s: Underwater medium shot. Kai glides down past the camera in a slow kick, left arm extended, the watch glinting as light caustics pass over the bracelet. Tiny bubbles trail from his nose. His body stays the same size and shape throughout. 5.0s to 7.5s: Underwater close-up. Kai's right hand turns the green bezel two clicks; the luminous markers glow softly in the deeper blue. His fingers grip the bezel edge naturally, five fingers visible. 7.5s to 10.0s: On deck, medium close-up. Kai climbs over the rail, pushes wet hair back, glances at the watch on his wrist, and says with a relaxed grin, lips synced: "Right on time." Sun flares gently behind him. SOUND: Deck: soft lap of waves against the hull, a gull far off, water dripping onto teak, and the tick of the watch captured close in the macro shot. Underwater: a muffled, enveloping hush with slow bubble sounds and the crisp mechanical click of the bezel, amplified slightly. Back on deck: wind, the boat creaking, and the line spoken clearly. A low, warm synth pad underneath the whole piece, never louder than the effects. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A gritty documentary-style moment on a commercial crab boat in the Bering Sea, four shots of hard physical work in bad weather, built to show heavy water, spray, and a consistent crew member across cuts. THE SUBJECT: Tanya, a white Alaskan deckhand in her forties with a wind-reddened face, pale blue eyes, a blond ponytail tucked under a black knit beanie, bright orange rubber rain bibs and jacket streaked with salt, thick orange rubber gloves, and black rubber boots. A second crew member, Ray, a Filipino American man in his fifties in matching orange gear with a gray mustache, appears once in the background operating the crane. SETTING: The steel deck of a crab fishing boat in heavy swells under a slate gray sky. A large steel-framed crab pot hangs from a crane, dripping seawater. Coiled blue lines, a sorting table, stacked pots, and towering dark waves beyond the rail. LIGHTING: Flat, cold overcast daylight with a hint of blue, sodium work lights on the mast adding a warm edge. Wet surfaces reflect specular highlights. Low contrast faces, high texture on gear, spray catching the light. CAMERA: Documentary handheld on a wide to normal lens, rolling with the boat, occasional water droplets striking the lens and running off. No slow motion. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Wide shot from the wheelhouse door. A big wave breaks over the bow and washes across the deck; Tanya braces with one hand on the rail, knees bent, as white water swirls around her boots. 2.5s to 5.0s: Medium shot. The crab pot swings in on the crane line, streaming water. Ray is visible behind at the crane controls. Tanya grabs the steel frame with both gloved hands and steadies it onto the sorting table with a heavy thud. 5.0s to 7.5s: Close-up over the table. The pot door opens and dozens of dark red king crabs spill across the steel, legs moving. Tanya scoops two and tosses them into a tank chute. 7.5s to 10.0s: Close-up on Tanya, spray on her cheeks, breathing hard. She looks past the camera toward the wheelhouse and shouts over the wind, lips synced: "Full pot! Send the next one!" Another wave rises behind her. SOUND: Roaring wind and crashing waves throughout, the diesel engine thrum under everything, the hydraulic whine of the crane, rope creaking, seawater pouring off steel, the heavy metallic slam of the pot on the table, crab shells clattering. Tanya's shout is loud but partly battered by wind. No music. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A music performance moment in a New York City subway station, four shots that build from a lonely player to a small audience, built to show accurate bowing and fingering synced to the sound. THE SUBJECT: Theo, a Korean American violinist in his mid twenties with a black undercut, round wire glasses, a charcoal wool overcoat over a black turtleneck, and a well-worn honey-colored violin with a dark chin rest. His open violin case lies at his feet lined with red velvet, a few dollar bills inside. A commuter, Denise, an African American nurse in her forties in navy scrubs and a puffy gray coat with a tote bag, stops to listen. SETTING: A tiled subway platform in Manhattan late on a winter evening. White tile walls, green steel columns, a yellow safety strip along the platform edge, a bench, and the dark tunnel mouth. A few scattered commuters in coats. LIGHTING: Overhead fluorescent fixtures give a cool white base; a warm vending machine glow on one side of Theo's face. Train headlights in the last shot throw a moving white glare down the tunnel. CAMERA: Intimate cinema style with a 50mm lens, slow push-ins and one gentle arc. Shallow depth of field so the tiles fall into soft bokeh. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 3.0s: Wide shot down the platform. Theo plays alone near a green column, eyes closed, bow moving in long smooth strokes. Two commuters walk past without stopping. 3.0s to 5.5s: Extreme close-up on the fingerboard. His left-hand fingers press and vibrate on the strings in perfect time with the melody; rosin dust lifts off the strings under the bow. 5.5s to 8.0s: Medium shot on Denise. She slows, stops beside the bench, lets her tote bag slide down her arm, and smiles. She says quietly to herself, lips synced: "Now that is how you end a shift." 8.0s to 10.0s: Slow arc around Theo as he finishes the phrase with a long held note. A train's headlights sweep in from the tunnel behind him and the platform fills with wind; his coat hem lifts. SOUND: A solo violin plays a slow, warm, original melody in a minor key, the bowing and fingering matching the picture precisely, with the natural echo of a tiled station. Faint footsteps and a distant PA chime. Denise's line is soft and close. In the final shot the rumble and screech of an arriving train rises and nearly swallows the last note. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A science fiction drama beat inside a research habitat on Mars, four widescreen shots with two characters, built to show consistent costumes, detailed production design, and a clean cinematic cut pattern. THE SUBJECT: Commander Ana Ruiz, a Puerto Rican astronaut in her late forties with close-cropped black hair going gray at the temples, a small scar through her left eyebrow, and a rust-stained white EVA suit with orange shoulder panels, helmet under her arm. Engineer Sam Okafor, a Nigerian American man in his early thirties with a shaved head and a short beard, wearing a charcoal habitat jumpsuit with rolled sleeves and a tablet strapped to his forearm. SETTING: The interior of a cramped Mars habitat module: curved white composite walls, exposed cable runs, a round airlock hatch with a red and green indicator ring, a small porthole showing a dusty red landscape and a pale butterscotch sky, and a bench with hanging EVA gear. LIGHTING: Cool white LED panels overhead, a red warning glow pulsing from the airlock ring before switching to green, and dusty daylight through the porthole. Moody contrast with clean practical sources. CAMERA: Anamorphic widescreen feel with gentle horizontal lens flares from practicals, slow dolly moves, a locked two-shot, and controlled close-ups. Filmic color with teal shadows and warm highlights. FRAMING (16:9 landscape): Compose for a wide screen. Use the full width: place the main subject on a rule-of-thirds vertical and let the setting breathe on the opposite side so the environment tells part of the story. Wide shots should read as true establishing shots with clear foreground, midground, and background layers. Close-ups can sit off center with negative space in the direction the character is looking. Keep the horizon level unless a shot below says otherwise. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Wide shot of the module. The airlock ring glows red, then turns green with a hiss; the hatch swings open and fine red dust drifts in around Ana as she steps through, helmet under her arm. 2.5s to 5.0s: Medium shot on Sam at the bench, he looks up from his forearm tablet, relief across his face, and says, lips synced: "You were gone for six hours." 5.0s to 7.5s: Close-up on Ana, dust on her cheek, a tired half smile. She holds up a small sealed sample tube with a pale blue crystal inside and answers, lips synced: "Worth every minute." 7.5s to 10.0s: Over-the-shoulder shot from behind Sam. Ana places the sample tube in his open palm; he turns it toward the porthole light, and the crystal throws a faint blue glint across both of their faces. SOUND: A low constant hum of life support fans and a soft electronic beep pattern. The airlock cycles with a pneumatic hiss and a heavy mechanical clunk. Boots on metal grating, suit fabric rustling. Dialogue is intimate with a slight metallic room reflection. A faint, sparse ambient score of low strings enters under the last shot. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A warm family cooking moment in a kitchen during the holidays, four shots with a grandmother teaching her grandson, built to show hands working dough precisely and a consistent pair of characters. THE SUBJECT: Abuela Rosa, a Mexican American grandmother in her seventies with silver hair in a low bun, gold hoop earrings, a lilac cardigan over a floral blouse, and a white apron dusted with masa. Her grandson Mateo, about ten years old, with messy black hair, a red hoodie with sleeves pushed up, and a serious, focused expression, the tip of his tongue poking out while he concentrates. Rosa's hands are small, wrinkled, and quick, with a thin gold wedding band; Mateo's are smooth and slightly sticky with masa. Both keep the same clothes and hair in every shot. SETTING: A small, lived-in family kitchen in San Antonio in December. A table covered with a checked oilcloth, a big metal bowl of masa, a stack of soaked corn husks, a pot of red pork filling, a large steamer pot on the stove, string lights in the window, and family photos on the fridge. LIGHTING: Soft window daylight from one side, warm tungsten from a ceiling fixture, and steam catching the light. Gentle and cozy with warm skin tones. CAMERA: Top-down and eye-level shots on a 35mm lens, steady handheld, square composition centered on the hands. FRAMING (1:1 square): Compose for a square feed post. Center the subject with even margins on all sides, and build each frame around one clear shape or gesture that reads at thumbnail size. Avoid important detail near the corners, since feeds and grids often round or crop them. Keep the background simple enough that the subject separates cleanly from it, using depth of field, a contrasting color, or a clean edge of light. Movement inside the frame should travel toward or away from the camera, or across the center, rather than off the sides. Medium shots and close-ups work best; wide shots should keep the subject large enough to recognize on a small screen. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Top-down close-up. Rosa's hands spread masa thinly across a husk with the back of a spoon in one smooth stroke. Mateo's smaller hands copy her on his own husk beside hers, clumsier and thicker. 2.5s to 5.0s: Eye-level medium two-shot. Rosa leans in and taps his husk with one finger, smiling, and says, lips synced: "Thinner, mijo. Like a whisper." 5.0s to 7.5s: Close-up on Mateo's hands. He scrapes the extra masa away, spoons in red filling, and folds the husk neatly into a packet, pressing the fold with his thumb. 7.5s to 10.0s: Medium shot. He holds up the finished tamale proudly. Rosa laughs, pulls him into a side hug, and kisses the top of his head as steam rolls up from the pot behind them. SOUND: Kitchen ambience: the simmer of the steamer pot, spoon scraping masa, husks rustling, a television playing a holiday program faintly in the next room, and relatives laughing far off. Rosa's line is warm and close. A soft acoustic guitar melody enters in the final shot. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A quiet short-film moment in a roadside diner after midnight, told in four shots with one exchange of dialogue, built to show off a single continuous performance carried across cuts. THE SUBJECT: Loretta, a Black American waitress in her late fifties with short silver-streaked natural hair, reading glasses pushed up on her head, a pale mint diner uniform dress with a white collar, a name tag with no readable text, and a small gold cross necklace. She moves with the ease of someone who has worked this counter for thirty years. Across the counter sits Danny, a white trucker in his early thirties with a sunburned neck, a sandy three-day beard, a faded red flannel shirt over a gray T-shirt, and a sweat-stained tan cap he takes off and holds in both hands. Props that stay consistent: a chipped white ceramic mug, a glass coffee pot two thirds full, and a slice of cherry pie on a plain white plate with a single fork. SETTING: A small 1970s roadside diner off an interstate in rural Nebraska at 1 a.m. Red vinyl stools, a long speckled Formica counter, chrome napkin dispensers, a pie case with a glowing lid, and a wide front window showing an empty highway, one blinking yellow traffic light, and a parked semi truck with its running lights on. Rain has just stopped; the window glass carries fine droplets. LIGHTING: Warm fluorescent tubes over the counter, slightly green in the shadows, mixed with the cool blue spill of the night outside the window and the amber pulse of the traffic light. Soft practical glow from the pie case lights the lower half of faces. Gentle contrast, deep but readable shadows, a subtle film grain. CAMERA: Shot on a 35mm cinema camera with vintage spherical lenses, shallow depth of field, slow deliberate moves. Mostly locked-off or slow push-ins at counter height. Natural color grade leaning warm, with the blacks lifted slightly like 35mm print film. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Medium shot from behind the counter at shoulder height. Loretta tops off the chipped mug in front of Danny; coffee streams in a steady arc and steam curls up. Danny sits with his cap in his hands, eyes on the pie. 2.5s to 5.0s: Close-up on Danny, slightly low angle, window and blinking light soft behind him. He looks up at her and speaks, voice tired and a little embarrassed: "I cannot pay for the pie." His lips sync to every word; he gives a small apologetic shrug. 5.0s to 7.5s: Close-up on Loretta, reverse angle matching his eyeline. A slow smile starts at the corners of her eyes. She sets the coffee pot down on the counter with a soft clink and answers warmly: "Pie is on me, honey." Her lips sync precisely. 7.5s to 10.0s: Wide shot from the far end of the counter, the whole diner visible. Danny lets out a breath, sets his cap on the stool beside him, and picks up the fork. Loretta walks away along the counter toward the kitchen pass, wiping her hands on a towel. The semi truck idles outside the window. SOUND: Continuous diner room tone: the hum of fluorescent tubes, a refrigerator compressor in the pie case, a distant truck passing on wet asphalt outside. Coffee pouring into the mug, the ceramic clink of the pot on Formica, the scrape of a fork on a plate at the end. A small radio in the kitchen plays an indistinct old country song very low, no clear lyrics. Dialogue is dry and close, with the room tone steady underneath and no music swell. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A tense sports moment at a small-town rodeo, four shots that move from calm preparation to explosive release, built to show how the model holds a rider and horse consistent through fast action. THE SUBJECT: Maya, a Mexican American barrel racer in her early twenties with a long dark braid down her back, a cream felt cowboy hat, a turquoise long-sleeve western shirt with pearl snaps, dark blue jeans, brown leather chaps, and worn brown boots. Her horse is a sorrel quarter horse mare named Pepper with a white blaze down her face, a black leather bridle with a silver concho, and a turquoise saddle pad that matches the shirt. The same horse and tack appear in every shot. SETTING: An outdoor county fair rodeo arena in west Texas at golden hour. A fenced dirt arena with three bright blue barrels, metal bleachers with a small crowd in the far background, the alley gate where riders wait, dust hanging in the air, and flat ranch land with a single windmill on the horizon. LIGHTING: Low late-afternoon sun from behind the rider, creating a warm rim light on the hat brim, the braid, and the horse mane. Dust in the air catches the light and glows. Faces are filled by soft bounce from the pale dirt. Rich warm tones, golden skin, deep blue sky. CAMERA: Sports documentary style. The first shots are handheld with a long lens and shallow depth of field; the release is a fast tracking shot. Slight natural camera shake, no stabilization wobble artifacts. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Tight close-up in the alley. Maya leans forward and presses her forehead to the mare's neck, eyes closed. She whispers softly, lips moving in sync: "Easy, girl. Just like home." Pepper flicks an ear back toward her voice. 2.5s to 4.5s: Extreme close-up of Maya's gloved hand shortening the leather reins, knuckles tightening, then her boot heel settling into the stirrup. Pepper paws the dirt once and a small puff of dust rises. 4.5s to 7.5s: Low wide tracking shot beside the arena fence. Pepper bursts out of the alley at full gallop toward the first blue barrel, dirt kicked high behind her hooves, Maya low over the saddle horn, braid flying. 7.5s to 10.0s: Medium shot as Pepper turns tight around the first barrel, leaning hard, Maya's inside hand on the horn, the barrel rocking but not falling. Dirt sprays toward the camera. The crowd in the far bleachers rises. SOUND: Alley shot: close, intimate sound of the horse breathing and snorting, leather creaking, the jingle of the bit, Maya's whisper. Then hoofbeats thunder in rhythm as the mare launches, dirt spraying, the rattle of the barrel as she turns, and a rising cheer from a small crowd with a few whistles. A distant arena announcer's voice is muffled and unintelligible. No music. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A science fiction drama beat inside a research habitat on Mars, four widescreen shots with two characters, built to show consistent costumes, detailed production design, and a clean cinematic cut pattern. THE SUBJECT: Commander Ana Ruiz, a Puerto Rican astronaut in her late forties with close-cropped black hair going gray at the temples, a small scar through her left eyebrow, and a rust-stained white EVA suit with orange shoulder panels, helmet under her arm. Engineer Sam Okafor, a Nigerian American man in his early thirties with a shaved head and a short beard, wearing a charcoal habitat jumpsuit with rolled sleeves and a tablet strapped to his forearm. SETTING: The interior of a cramped Mars habitat module: curved white composite walls, exposed cable runs, a round airlock hatch with a red and green indicator ring, a small porthole showing a dusty red landscape and a pale butterscotch sky, and a bench with hanging EVA gear. LIGHTING: Cool white LED panels overhead, a red warning glow pulsing from the airlock ring before switching to green, and dusty daylight through the porthole. Moody contrast with clean practical sources. CAMERA: Anamorphic widescreen feel with gentle horizontal lens flares from practicals, slow dolly moves, a locked two-shot, and controlled close-ups. Filmic color with teal shadows and warm highlights. FRAMING (16:9 landscape): Compose for a wide screen. Use the full width: place the main subject on a rule-of-thirds vertical and let the setting breathe on the opposite side so the environment tells part of the story. Wide shots should read as true establishing shots with clear foreground, midground, and background layers. Close-ups can sit off center with negative space in the direction the character is looking. Keep the horizon level unless a shot below says otherwise. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Wide shot of the module. The airlock ring glows red, then turns green with a hiss; the hatch swings open and fine red dust drifts in around Ana as she steps through, helmet under her arm. 2.5s to 5.0s: Medium shot on Sam at the bench, he looks up from his forearm tablet, relief across his face, and says, lips synced: "You were gone for six hours." 5.0s to 7.5s: Close-up on Ana, dust on her cheek, a tired half smile. She holds up a small sealed sample tube with a pale blue crystal inside and answers, lips synced: "Worth every minute." 7.5s to 10.0s: Over-the-shoulder shot from behind Sam. Ana places the sample tube in his open palm; he turns it toward the porthole light, and the crystal throws a faint blue glint across both of their faces. SOUND: A low constant hum of life support fans and a soft electronic beep pattern. The airlock cycles with a pneumatic hiss and a heavy mechanical clunk. Boots on metal grating, suit fabric rustling. Dialogue is intimate with a slight metallic room reflection. A faint, sparse ambient score of low strings enters under the last shot. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A premium product film for a stainless steel dive watch, four shots that alternate macro detail with a person wearing it underwater and on a boat, built for a vertical ad with one short spoken line. THE SUBJECT: The product: a stainless steel dive watch with a 42mm brushed case, a deep green sunburst dial, a rotating ceramic bezel in matching green with luminous markers, white luminous hands, and a brushed steel bracelet. No brand name or text is visible on the dial. The wearer is Kai, a Native Hawaiian free diver in his thirties with short black hair, a trimmed beard, a sun-darkened complexion, a simple black wetsuit top unzipped to the chest, and the watch on his left wrist. SETTING: Clear turquoise water off the coast of Maui over a white sand bottom with scattered lava rock, then the teak deck of a small white dive boat, wet from spray, with coiled rope and a folded mask resting on the rail. LIGHTING: Underwater: sunlight cuts down in moving caustic patterns across the sand and across the watch, blue-green color, soft light shafts. On deck: bright late-morning sun with crisp specular highlights sliding across the brushed steel and the domed crystal, deep blue sky behind. CAMERA: High-end commercial look. Macro probe lens for the product, slow motorized sliders, and a stabilized underwater housing. Crisp focus on the dial, gentle focus falloff everywhere else, luxurious slow motion feel without actual frame skipping. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Macro shot on deck. The watch rests on wet teak; a single water droplet rolls across the domed crystal while the camera glides left, a highlight sweeping across the green sunburst dial. 2.5s to 5.0s: Underwater medium shot. Kai glides down past the camera in a slow kick, left arm extended, the watch glinting as light caustics pass over the bracelet. Tiny bubbles trail from his nose. His body stays the same size and shape throughout. 5.0s to 7.5s: Underwater close-up. Kai's right hand turns the green bezel two clicks; the luminous markers glow softly in the deeper blue. His fingers grip the bezel edge naturally, five fingers visible. 7.5s to 10.0s: On deck, medium close-up. Kai climbs over the rail, pushes wet hair back, glances at the watch on his wrist, and says with a relaxed grin, lips synced: "Right on time." Sun flares gently behind him. SOUND: Deck: soft lap of waves against the hull, a gull far off, water dripping onto teak, and the tick of the watch captured close in the macro shot. Underwater: a muffled, enveloping hush with slow bubble sounds and the crisp mechanical click of the bezel, amplified slightly. Back on deck: wind, the boat creaking, and the line spoken clearly. A low, warm synth pad underneath the whole piece, never louder than the effects. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A warm family cooking moment in a kitchen during the holidays, four shots with a grandmother teaching her grandson, built to show hands working dough precisely and a consistent pair of characters. THE SUBJECT: Abuela Rosa, a Mexican American grandmother in her seventies with silver hair in a low bun, gold hoop earrings, a lilac cardigan over a floral blouse, and a white apron dusted with masa. Her grandson Mateo, about ten years old, with messy black hair, a red hoodie with sleeves pushed up, and a serious, focused expression, the tip of his tongue poking out while he concentrates. Rosa's hands are small, wrinkled, and quick, with a thin gold wedding band; Mateo's are smooth and slightly sticky with masa. Both keep the same clothes and hair in every shot. SETTING: A small, lived-in family kitchen in San Antonio in December. A table covered with a checked oilcloth, a big metal bowl of masa, a stack of soaked corn husks, a pot of red pork filling, a large steamer pot on the stove, string lights in the window, and family photos on the fridge. LIGHTING: Soft window daylight from one side, warm tungsten from a ceiling fixture, and steam catching the light. Gentle and cozy with warm skin tones. CAMERA: Top-down and eye-level shots on a 35mm lens, steady handheld, square composition centered on the hands. FRAMING (1:1 square): Compose for a square feed post. Center the subject with even margins on all sides, and build each frame around one clear shape or gesture that reads at thumbnail size. Avoid important detail near the corners, since feeds and grids often round or crop them. Keep the background simple enough that the subject separates cleanly from it, using depth of field, a contrasting color, or a clean edge of light. Movement inside the frame should travel toward or away from the camera, or across the center, rather than off the sides. Medium shots and close-ups work best; wide shots should keep the subject large enough to recognize on a small screen. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Top-down close-up. Rosa's hands spread masa thinly across a husk with the back of a spoon in one smooth stroke. Mateo's smaller hands copy her on his own husk beside hers, clumsier and thicker. 2.5s to 5.0s: Eye-level medium two-shot. Rosa leans in and taps his husk with one finger, smiling, and says, lips synced: "Thinner, mijo. Like a whisper." 5.0s to 7.5s: Close-up on Mateo's hands. He scrapes the extra masa away, spoons in red filling, and folds the husk neatly into a packet, pressing the fold with his thumb. 7.5s to 10.0s: Medium shot. He holds up the finished tamale proudly. Rosa laughs, pulls him into a side hug, and kisses the top of his head as steam rolls up from the pot behind them. SOUND: Kitchen ambience: the simmer of the steamer pot, spoon scraping masa, husks rustling, a television playing a holiday program faintly in the next room, and relatives laughing far off. Rosa's line is warm and close. A soft acoustic guitar melody enters in the final shot. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A gritty documentary-style moment on a commercial crab boat in the Bering Sea, four shots of hard physical work in bad weather, built to show heavy water, spray, and a consistent crew member across cuts. THE SUBJECT: Tanya, a white Alaskan deckhand in her forties with a wind-reddened face, pale blue eyes, a blond ponytail tucked under a black knit beanie, bright orange rubber rain bibs and jacket streaked with salt, thick orange rubber gloves, and black rubber boots. A second crew member, Ray, a Filipino American man in his fifties in matching orange gear with a gray mustache, appears once in the background operating the crane. SETTING: The steel deck of a crab fishing boat in heavy swells under a slate gray sky. A large steel-framed crab pot hangs from a crane, dripping seawater. Coiled blue lines, a sorting table, stacked pots, and towering dark waves beyond the rail. LIGHTING: Flat, cold overcast daylight with a hint of blue, sodium work lights on the mast adding a warm edge. Wet surfaces reflect specular highlights. Low contrast faces, high texture on gear, spray catching the light. CAMERA: Documentary handheld on a wide to normal lens, rolling with the boat, occasional water droplets striking the lens and running off. No slow motion. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Wide shot from the wheelhouse door. A big wave breaks over the bow and washes across the deck; Tanya braces with one hand on the rail, knees bent, as white water swirls around her boots. 2.5s to 5.0s: Medium shot. The crab pot swings in on the crane line, streaming water. Ray is visible behind at the crane controls. Tanya grabs the steel frame with both gloved hands and steadies it onto the sorting table with a heavy thud. 5.0s to 7.5s: Close-up over the table. The pot door opens and dozens of dark red king crabs spill across the steel, legs moving. Tanya scoops two and tosses them into a tank chute. 7.5s to 10.0s: Close-up on Tanya, spray on her cheeks, breathing hard. She looks past the camera toward the wheelhouse and shouts over the wind, lips synced: "Full pot! Send the next one!" Another wave rises behind her. SOUND: Roaring wind and crashing waves throughout, the diesel engine thrum under everything, the hydraulic whine of the crane, rope creaking, seawater pouring off steel, the heavy metallic slam of the pot on the table, crab shells clattering. Tanya's shout is loud but partly battered by wind. No music. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Prompt
A music performance moment in a New York City subway station, four shots that build from a lonely player to a small audience, built to show accurate bowing and fingering synced to the sound. THE SUBJECT: Theo, a Korean American violinist in his mid twenties with a black undercut, round wire glasses, a charcoal wool overcoat over a black turtleneck, and a well-worn honey-colored violin with a dark chin rest. His open violin case lies at his feet lined with red velvet, a few dollar bills inside. A commuter, Denise, an African American nurse in her forties in navy scrubs and a puffy gray coat with a tote bag, stops to listen. SETTING: A tiled subway platform in Manhattan late on a winter evening. White tile walls, green steel columns, a yellow safety strip along the platform edge, a bench, and the dark tunnel mouth. A few scattered commuters in coats. LIGHTING: Overhead fluorescent fixtures give a cool white base; a warm vending machine glow on one side of Theo's face. Train headlights in the last shot throw a moving white glare down the tunnel. CAMERA: Intimate cinema style with a 50mm lens, slow push-ins and one gentle arc. Shallow depth of field so the tiles fall into soft bokeh. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 3.0s: Wide shot down the platform. Theo plays alone near a green column, eyes closed, bow moving in long smooth strokes. Two commuters walk past without stopping. 3.0s to 5.5s: Extreme close-up on the fingerboard. His left-hand fingers press and vibrate on the strings in perfect time with the melody; rosin dust lifts off the strings under the bow. 5.5s to 8.0s: Medium shot on Denise. She slows, stops beside the bench, lets her tote bag slide down her arm, and smiles. She says quietly to herself, lips synced: "Now that is how you end a shift." 8.0s to 10.0s: Slow arc around Theo as he finishes the phrase with a long held note. A train's headlights sweep in from the tunnel behind him and the platform fills with wind; his coat hem lifts. SOUND: A solo violin plays a slow, warm, original melody in a minor key, the bowing and fingering matching the picture precisely, with the natural echo of a tiled station. Faint footsteps and a distant PA chime. Denise's line is soft and close. In the final shot the rumble and screech of an arriving train rises and nearly swallows the last note. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.
Why creators choose Seedance 2.5
Clips up to 30 seconds
Seedance 2.5 generates anywhere from 4 to 30 seconds in a single take, twice the 15 second ceiling of the Seedance 2.0, 2.0 Fast, and 2.0 Mini tiers. That is enough runway for a full ad beat or a short scene.
Several shots in one take
ByteDance built Seedance 2.5 to organize multiple connected shots inside one generation, with setup, development, and payoff. Write timecoded cuts in your prompt and the model follows them.
Audio and video generated together
Sound is produced in the same pass as the picture: dialogue with lip sync, sound effects, and room tone that land on the action. No separate audio step before you publish.
The widest reference input in the family
Guide a generation with up to 30 images, 10 video clips, and 10 audio clips at once on Fliki. Seedance 2.0 and its Fast and Mini tiers top out at 9 images, 3 videos, and 3 audio clips.
Consistent characters and props
Faces, wardrobe, and product details hold across cuts, so a character introduced in shot one is still recognizably the same person in shot four.
Follows long, detailed prompts
Seedance 2.5 accepts prompts up to 10,000 characters on Fliki and reads them closely. Camera language, blocking, and per-shot timing all come through.
Believable motion and physics
Cloth, liquids, and bodies move with weight. Scenes feel filmed rather than animated, which matters most for lifestyle, product, and narrative work.
Native 16:9, 9:16, and 1:1
Compose for YouTube, Shorts and Reels, or feed posts from the same prompt. Each ratio is framed natively, not cropped from another.
How it works
How to generate a video with Seedance 2.5
Seedance 2.5 runs in the Fliki AI Playground. Follow these six steps to get your first clip.

Write your prompt
Open Fliki and describe the scene in plain language: subject, setting, lighting, camera, and any dialogue in quotes. Seedance 2.5 follows long, detailed prompts, so write it like a shot list with timecodes if you want cuts.

Select Seedance 2.5 as your model
Open the AI Playground in Fliki and choose Seedance 2.5 from the video model list. Fliki sends your prompt straight to ByteDance's model with no extra setup.

Pick your aspect ratio
Choose 16:9 for YouTube and landscape web, 9:16 for TikTok, Reels, and Shorts, or 1:1 for feed posts. Seedance 2.5 composes each ratio natively.

Set the duration
Pick any whole number of seconds from 4 to 30. Short clips are good for testing a look; longer ones give Seedance 2.5 room to play out several shots with a beginning, middle, and end.

Add references (optional)
Attach up to 30 reference images, 10 video clips, and 10 audio clips, plus optional first and last frames. Use them to lock a character, a product, a camera move, or a voice.

Select resolution and generate
Seedance 2.5 renders at 720p on Fliki. Hit Generate, then preview, download, or bring the clip into a longer Fliki project.
AI MODEL GALLERY
Built on the best AI models - ready inside Fliki
Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.
















Seedance 2.5 FAQ
Frequently asked questions
Everything you need to know about generating with Seedance 2.5 inside Fliki.
Seedance 2.5 is ByteDance Seed's video and audio generation model, announced on July 31, 2026. It generates up to 30 seconds per take, can plan several connected shots in one generation, and produces synchronized sound in the same pass.
Seedance 2.5 is available in Fliki's AI Playground. It is not offered in the regular video creation pickers, which use tiers like Seedance 2.0 Fast and Seedance 2.0 Mini instead.
Seedance 2.5 is available on the Premium plan. It is not included on the Free, Basic, or Standard plans.
On Fliki you can pick any whole number of seconds from 4 to 30. The Seedance 2.0 tiers stop at 15 seconds.
Fliki renders Seedance 2.5 at 720p in 16:9, 9:16, or 1:1.
Yes. Dialogue, sound effects, and ambience are generated together with the video, so lip movement and on-screen action line up with the sound.
On Fliki you can attach up to 30 reference images, 10 reference videos, and 10 reference audio clips, plus a first and last frame image.
Seedance 2.5 is the newest and most capable tier: longer clips (up to 30 seconds), far more reference inputs, and multi-shot planning built in. Seedance 2.0 is the previous flagship. Seedance 2.0 Fast and 2.0 Mini trade some quality for speed and a lower credit cost, stop at 15 seconds, and are available on more plans.
Still curious?
Try Fliki free in your browser, no credit card required.
Start free→More from Fliki
AI video models
Discover more
Tools
Discover features
Generate your next video with Seedance 2.5.
ByteDance's longest-running Seedance: multi-shot takes up to 30 seconds with sound generated alongside the picture. Free to start, no credit card required.
Generate your first video freeFree forever plan · No credit card required · Cancel anytime


