video model · by ByteDance

Seedance 2.0 Mini AI Video Generator

Seedance 2.0 Mini is the lightest, lowest-cost tier of ByteDance's Seedance 2.0 family. It keeps multi-shot cuts, native audio, and image, video, and audio references, renders at 480p or 720p, and is built for fast, affordable clips across Fliki's video tools. Compare it side by side with all our AI video models before you render.

Generated with Seedance 2.0 Mini

Seedance 2.0 Mini clips generated inside Fliki. No edits, no post.

Prompt

A friendly vertical clip for a pediatric dental office, four gentle shots with one reassuring line, built to show warm performances and a consistent child and dentist. THE SUBJECT: Dr. Grace Liu, a Chinese American pediatric dentist in her forties with a sleek black bob, tortoiseshell glasses, light blue scrubs under a white coat, and a patterned surgical cap with small cartoon stars. Her patient is Noah, a white boy around six with curly light brown hair, freckles, and a green dinosaur T-shirt, clutching a small plush triceratops. SETTING: A cheerful pediatric dental office in suburban Minneapolis with a mint green dental chair, a ceiling mural of clouds and hot air balloons, a tray with a small round mirror, and a window with afternoon light. LIGHTING: Soft, even daylight from the window, a warm overhead dental lamp on Noah's face, and gentle pastel tones. CAMERA: Steady eye-level shots on a 35mm lens with slow push-ins and a soft, friendly grade. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Medium shot. Noah sits stiffly in the big chair, hugging the plush triceratops, eyes wide. 2.5s to 5.0s: Close-up on Dr. Liu. She crouches to his eye level, holds up the little mirror, and says softly, lips synced: "Can we count your teeth together?" 5.0s to 7.5s: Close-up on Noah. He opens his mouth wide while holding the dinosaur up so it can see too. 7.5s to 10.0s: Medium two-shot. Dr. Liu gives him a high five and he beams, relaxed now, swinging his sneakers off the edge of the chair. SOUND: Quiet office ambience, a soft tick of the mirror on the tray, the chair creaking, a distant receptionist's muffled voice, and a gentle playful piano tune underneath. The high five lands with a light clap. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

Prompt

A slow, satisfying clip from a small family apiary, four shots from hive to jar, built to show thick golden honey, bees in motion, and a consistent beekeeper. THE SUBJECT: Walter, a white beekeeper in his sixties with a gray beard, a white ventilated bee jacket with the veil pushed back, tan work pants, and yellow nitrile gloves. Props: a wooden hive frame heavy with capped honeycomb, a stainless uncapping knife, and a glass mason jar. SETTING: A small apiary at the edge of a wildflower meadow in rural Vermont in late summer, with white wooden hive boxes, purple clover and goldenrod, and a weathered red barn in the soft background. LIGHTING: Warm late-afternoon sun backlighting the comb so the honey glows amber, bees lit in silhouette against the light. CAMERA: Macro and medium shots with a gentle handheld float, shallow depth of field, warm natural grade. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Medium shot. Walter lifts a frame from the hive; dozens of bees crawl across the golden comb. 2.5s to 5.0s: Macro close-up. The uncapping knife slides across the wax caps and honey wells up and runs down the comb. 5.0s to 7.5s: Close-up. A thick ribbon of honey pours from a spout into the mason jar, folding over itself. 7.5s to 10.0s: Medium close-up. Walter holds the full jar up to the sun, smiles, and says, lips synced: "This is the good stuff." SOUND: A rich steady hum of bees, crickets in the meadow, the soft scrape of the knife through wax, the slow glug of honey into glass, and a light breeze. No music. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

Prompt

A cozy vertical recommendation clip for an independent bookstore, four shots ending on a quick pitch, built to show warm interiors and a consistent bookseller. THE SUBJECT: Amara, a Black American bookseller in her early thirties with shoulder-length locs tied back with a mustard scarf, gold round glasses, a cream cable-knit cardigan, and a brown corduroy skirt. The prop is a hardcover novel with a deep blue dust jacket and no readable title. SETTING: A small independent bookstore in Brooklyn on an autumn evening, with tall wooden shelves, a rolling ladder, hand-lettered shelf signs with no readable text, a sleeping orange cat on a stack of books, and string lights in the window. LIGHTING: Warm tungsten lamps and string lights, deep amber tones, soft pools of light on the shelves and her face. CAMERA: Slow dolly moves along the shelves on a 50mm lens, shallow depth of field, cozy warm grade. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Slow dolly along a shelf; Amara's finger runs across the spines and stops on the blue book. 2.5s to 5.0s: Close-up as she pulls it out and flips it open; pages fan past her thumb. 5.0s to 7.5s: Medium close-up. She holds the book to her chest and says to camera, lips synced: "I finished it in one night." 7.5s to 10.0s: Wide shot. She sets the book on the front table beside the sleeping cat, who opens one eye and stretches. SOUND: Pages rustling, the soft slide of a book off the shelf, a wooden floor creak, rain outside, and a quiet jazz record playing in the shop. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

Prompt

A feel-good clip from a dog grooming salon, four shots from a scruffy arrival to a fluffy finish, built to show fur detail, a happy dog, and consistent characters. THE SUBJECT: Biscuit, an apricot goldendoodle with shaggy curly fur and a red collar with a round silver tag. The groomer is Carmen, a Dominican American woman in her thirties with curly dark hair in a high bun, a teal grooming smock, and small hoop earrings. SETTING: A bright grooming salon in Tampa with a stainless steel grooming table, a raised tub, a wall of colorful bandanas, and a window showing palm trees. LIGHTING: Clean bright daylight with soft overhead LEDs, fur highlighted by a gentle backlight. CAMERA: Steady handheld medium shots and close-ups on a 35mm lens, cheerful bright grade. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Medium shot in the tub. Warm water sprays over Biscuit's curls and suds pile up as Carmen scrubs behind his ears. 2.5s to 5.0s: Close-up on the table. A blow dryer puffs his fur into soft waves; his ears flap and he squints happily. 5.0s to 7.5s: Close-up. Carmen trims around his face with rounded scissors, then ties a yellow bandana around his neck. 7.5s to 10.0s: Medium shot. Biscuit shakes once, fluffy and clean, and Carmen laughs and says, lips synced: "Look at you, handsome!" SOUND: Running water and splashing, the roar of the dryer, snipping scissors, jingling tags, a happy bark, and Carmen's laugh. A light ukulele tune underneath. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

Prompt

A small romantic short-film beat at a city bus stop in the rain, four shots with one line, built to show consistent strangers, falling rain, and a simple emotional turn. THE SUBJECT: Elena, a Filipino American graphic designer in her late twenties with a chin-length black bob, a mustard raincoat, and a canvas portfolio case. Ben, a white man in his early thirties with messy brown hair, a navy peacoat, and a large clear bubble umbrella. SETTING: A glass bus shelter on a quiet Seattle street at dusk in heavy rain, puddles on the sidewalk, wet brick buildings, and warm shop windows across the street. LIGHTING: Blue dusk light mixed with warm shop window glow and the cool light of the shelter. Rain drops catch backlight. CAMERA: Cinematic 50mm lens, slow push-ins, shallow depth of field, soft moody grade. FRAMING (9:16 vertical): Compose every shot for a phone held upright. Keep faces and the key action in the middle vertical band, roughly between one third and two thirds of frame height, so nothing important sits under the top status area or the bottom caption and button zone of Reels, TikTok, and Shorts. Favor medium close-ups and tall compositions that use depth front to back rather than width. When two people share a frame, stack them in depth or stagger them, never side by side at the far edges. Headroom stays modest; the eyes of the speaking character sit about one third down from the top. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Wide shot. Elena runs into the shelter, holding the portfolio over her head, rain pouring off the roof edge. 2.5s to 5.0s: Medium shot. The bus pulls away without her; she sighs and looks at her soaked sleeves. 5.0s to 7.5s: Medium two-shot. Ben steps in beside her, tilts the clear umbrella, and says with a shy smile, lips synced: "Want to share this?" 7.5s to 10.0s: Wide shot from across the street. They walk off together under the clear umbrella, rain streaming around its edges. SOUND: Heavy rain drumming on the shelter roof, tires hissing through puddles, the bus doors thumping shut and the engine pulling away, and a soft piano melody entering in the last shot. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

Prompt

A flat illustrated explainer in a clean motion graphics style, four shots that follow sunlight into a home, built to show the kind of illustrated look Seedance 2.0 Mini handles well. STYLE: flat 2D vector illustration with bold shapes, no outlines, a limited palette of sky blue, sunny yellow, leaf green, warm orange, and off-white, soft long shadows, subtle paper grain, and smooth animated motion. Not photoreal. THE SUBJECT: A small illustrated family house with a dark blue solar panel array on its roof, a round smiling sun character with simple dot eyes, glowing yellow energy particles, and a friendly illustrated homeowner, Mr. Patel, a middle-aged Indian American man with a mustache, a green sweater, and a coffee mug. SETTING: A simple illustrated suburban street with rounded trees, a picket fence, and gentle hills under a clear blue sky, then a cutaway view inside the house showing a kitchen. LIGHTING: Flat bright illustration lighting with soft, long, slightly transparent shadows and a warm glow around the sun and energy particles. CAMERA: Smooth 2D camera moves: slow pans, a push into the roof, and a cutaway reveal. Clean and steady. FRAMING (16:9 landscape): Compose for a wide screen. Use the full width: place the main subject on a rule-of-thirds vertical and let the setting breathe on the opposite side so the environment tells part of the story. Wide shots should read as true establishing shots with clear foreground, midground, and background layers. Close-ups can sit off center with negative space in the direction the character is looking. Keep the horizon level unless a shot below says otherwise. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Wide shot of the street. The sun character rises over the hill and beams of light stretch down toward the house roof. 2.5s to 5.0s: Push in on the panels. Light beams hit them and turn into small glowing yellow particles that flow down a wire along the wall. 5.0s to 7.5s: Cutaway inside. The particles travel through a small wall box and light up the kitchen lamp and the coffee maker. 7.5s to 10.0s: Medium shot. Mr. Patel lifts his mug, smiles, and says, lips synced: "Coffee, powered by sunshine." SOUND: Playful light sound design: a soft chime as the sun rises, a sparkly whoosh as light hits the panels, a gentle electric hum as the particles flow, the click of a lamp and the gurgle of the coffee maker. Upbeat light pizzicato music underneath. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

Prompt

A satisfying square clip from a ceramics studio, four shots from a lump of clay to a finished bowl shape, built to show wet clay physics and accurate hands. THE SUBJECT: June, a Japanese American potter in her fifties with a silver-streaked black braid, a rust-colored linen apron over a gray long-sleeve shirt, sleeves pushed up, and forearms streaked with wet clay. SETTING: A small sunlit pottery studio in Santa Fe with an electric wheel, a bucket of water, a sponge, wooden tools, shelves of drying bowls, and adobe walls. LIGHTING: Soft warm sunlight from a high window, a glossy sheen on the wet clay, gentle shadows. CAMERA: Close and medium shots on a 50mm lens, locked-off and slow push-ins, centered for square. FRAMING (1:1 square): Compose for a square feed post. Center the subject with even margins on all sides, and build each frame around one clear shape or gesture that reads at thumbnail size. Avoid important detail near the corners, since feeds and grids often round or crop them. Keep the background simple enough that the subject separates cleanly from it, using depth of field, a contrasting color, or a clean edge of light. Movement inside the frame should travel toward or away from the camera, or across the center, rather than off the sides. Medium shots and close-ups work best; wide shots should keep the subject large enough to recognize on a small screen. SHOTS (four shots joined by three hard cuts, 10 seconds total): 0.0s to 2.5s: Close-up. June slaps a lump of terracotta clay onto the center of the spinning wheel and cups it with both wet hands. 2.5s to 5.0s: Close-up. Her thumbs press into the center and the clay opens into a wide ring, spinning smoothly. 5.0s to 7.5s: Close-up. Her fingers pull the walls up into a bowl shape, water and slip running between her knuckles. 7.5s to 10.0s: Medium shot. The wheel slows; she sits back, wipes her hands on the apron, and says, lips synced: "That one is a keeper, I think." SOUND: The steady whir of the wheel, wet squelching clay, water dripping into the bucket, a soft wet slap at the start, and a quiet acoustic guitar underneath. CONTINUITY AND PERFORMANCE: Treat the four shots as one continuous scene filmed with the same cast on the same day, so time of day, weather, wet or dry surfaces, steam, dust, and the position of every prop carry logically from one shot to the next. Screen direction stays consistent: a character who looks or moves left in one shot is still oriented the same way after the cut unless the camera deliberately moves around them. Performances are restrained and natural, never theatrical; small reactions in the eyes and breath matter more than big gestures. Each shot begins with the action already underway, with no frozen first frame and no slow fade in, and the final shot ends on a held beat rather than a sudden stop. DETAIL PRIORITIES: Keep each composition simple and bold so it reads well at 480p as well as 720p. The main subject stays large in frame and clearly separated from the background by light, color, or depth of field; backgrounds are uncluttered, with no tiny text, busy patterns, or crowds of small faces. Give each shot one clear action that begins and ends inside its time slot, so every cut lands on a fresh, readable beat. Faces in close-up stay stable with natural blinking and natural teeth when talking or smiling. Colors stay consistent across all four shots as if graded together, with clean whites and no drifting tints. Favor steady or slow camera moves over fast whip pans so fine detail does not smear. Hands stay clean and well formed, with fingers clearly separated whenever they hold, pour, trim, or shape something close to the camera. RULES: Photorealistic live-action footage unless the style section above says otherwise, with natural skin texture, pores, fine hair, and believable fabric weight. Every character keeps the exact same face, hairstyle, wardrobe, and accessories in every shot; props keep the same color, size, labels, and position unless an action in the shots moves them. Hands are anatomically correct with five fingers each, natural knuckles, and a firm, believable grip on anything they hold; no fingers merging into objects. Eyelines match across cuts so conversations and reactions read correctly. Motion obeys real physics: liquids pour and splash with weight, cloth swings and settles, footsteps land. Cuts are clean hard cuts exactly at the listed timecodes, with no dissolves, no morphing between shots, and no warping of faces or backgrounds during camera moves. The total running time is exactly 10 seconds. No on-screen text, no captions, no subtitles, no logos, no watermarks, no brand names, no UI overlays, no letterbox bars, no split screens. Spoken lines are delivered exactly as written in quotes, in natural American English at a relaxed conversational pace, with mouth shapes precisely lip-synced to every syllable; only the character named for a line moves their lips while speaking it. No narrator, no voiceover, and no extra words beyond the quoted dialogue.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose Seedance 2.0 Mini

Lowest cost in the Seedance family

Seedance 2.0 Mini uses the fewest credits per second of any Seedance tier on Fliki, so you can generate many variations without burning through your plan.

Quick turnaround

Mini is built for speed. It suits drafts, social posts, and batch work where getting many usable clips matters more than peak detail.

Native audio

Mini generates a synchronized audio track with the video, including dialogue, effects, and ambience, so a clip is ready to post.

Multi-shot cuts

Timecoded shot lists come back as clean hard cuts inside one clip, with the same characters and props carried through.

Image, video, and audio references

Use up to 9 reference images, 3 reference videos, and 3 audio clips on Fliki, or a first frame image, to steer the result.

Strong on illustrated styles

Flat illustration, motion graphics, and explainer looks hold up well at 480p, which is why Fliki uses Mini as a default renderer for explainer videos.

Follows detailed prompts

Mini accepts prompts up to 10,000 characters on Fliki and handles camera language and per-shot timing faithfully.

Native 16:9, 9:16, and 1:1

Frame for YouTube, Shorts and Reels, or feed posts from one prompt, composed natively rather than cropped.

How it works

How to generate a video with Seedance 2.0 Mini

Seedance 2.0 Mini is the quickest way to get a Seedance clip inside Fliki. Follow these six steps.

Fliki prompt input with a shot-by-shot description for Seedance 2.0 Mini
Step 1

Write your prompt

Open Fliki and describe the scene in plain language: subject, setting, lighting, camera, and any dialogue in quotes. Seedance 2.0 Mini follows long, detailed prompts, so write it like a shot list with timecodes if you want cuts.

Fliki model selector with Seedance 2.0 Mini chosen
Step 2

Select Seedance 2.0 Mini as your model

Open the model selector in Fliki and choose Seedance 2.0 Mini. It is available in the regular video tools as well as the AI Playground.

Choose 16:9, 9:16, or 1:1 aspect ratio for Seedance 2.0 Mini on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube and landscape web, 9:16 for TikTok, Reels, and Shorts, or 1:1 for feed posts. Seedance 2.0 Mini composes each ratio natively.

Set clip duration on the Fliki slider for Seedance 2.0 Mini
Step 4

Set the duration

Pick any whole number of seconds from 4 to 15. A 10 second clip comfortably holds three or four short shots.

Upload reference images, video, or audio for Seedance 2.0 Mini on Fliki
Step 5

Add references (optional)

Attach up to 9 reference images, 3 video clips, and 3 audio clips, plus optional first and last frames, to keep a character or product on model.

Pick output resolution and generate with Seedance 2.0 Mini on Fliki
Step 6

Select resolution and generate

Choose 480p for the lowest cost, which suits illustrated and simple scenes well, or 720p for live-action looks. Then hit Generate.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

Seedance 2.0 Mini FAQ

Frequently asked questions

Everything you need to know about generating with Seedance 2.0 Mini inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
Seedance 2.0 Mini · Free forever plan

Generate your next video with Seedance 2.0 Mini.

The lowest-cost Seedance tier, with sound, multi-shot cuts, and reference input. Free to start, no credit card required.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime