video model · by MiniMax
MiniMax H3 Max Turbo AI Video Generator
Render 768p or 480p videos quickly with MiniMax H3 Max Turbo, the distilled speed variant of H3 Max. It keeps strong prompt adherence and native sound at about half the per-second price of H3 Max, with first and last frame guidance and clips up to 15 seconds. Compare it side by side with all our AI video models before you render.
Generated with MiniMax H3 Max Turbo
A handful of MiniMax H3 Max Turbo clips generated inside Fliki. No edits, no post.
Prompt
A 10-second 9:16 personal finance tip video filmed in a home office in Charlotte, North Carolina, where a creator explains the 24-hour rule for purchases. THE SUBJECT: Alicia, a 33-year-old Black American creator with a sleek middle-part bob, a mustard sweater, and gold hoops. SETTING: A home office with a plant, a bookshelf, and a laptop. LIGHTING: Soft window light with a warm lamp. CAMERA: Talking head medium with inserts. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: Alicia says: "Want it? Wait a day." 2.5s to 5.0s: Insert: she closes a laptop with a full online cart. 5.0s to 7.5s: Medium: she says: "Most urges fade overnight." 7.5s to 10.0s: Close-up: she smiles and taps a jar of coins. DETAILS: The home office has a white desk, a pothos plant trailing off the shelf, a stack of colorful books, a ceramic mug of pens, and a clear glass jar half full of coins with a handwritten paper label on it. The laptop screen in the insert shows a generic online shopping cart page with product thumbnails and no readable brand names or text. Alicia sits centered in shots one and three, framed from mid-chest up, looking straight into the lens with confident, friendly energy. She uses one small hand gesture on each line: a raised finger on the first line and an open palm on the second. Her sweater, bob, and hoops stay identical. In shot four she leans in, taps the jar lightly, and the coins shift. The pacing feels like a crisp educational short: calm, clear, and trustworthy, with a warm smile. DIALOGUE AND LIP-SYNC: Alicia speaks in shots one and three. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Quiet room, laptop click, coins clinking. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Clean explainer. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 cooking tutorial in a home kitchen in Seattle, Washington, where a chef shows the claw grip while dicing an onion. THE SUBJECT: Chef Min-jun, a 41-year-old Korean American man with short black hair and a white chef jacket, holding a chef knife. SETTING: A wooden cutting board, a yellow onion, and a tiled backsplash. LIGHTING: Bright soft overhead light. CAMERA: Top-down and medium. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: he says: "Tuck your fingertips in." 2.5s to 5.0s: Top-down: claw grip slicing the onion. 5.0s to 7.5s: Top-down: quick even dice. 7.5s to 10.0s: Medium: he says: "Safe and fast." DETAILS: The chef knife is an eight-inch steel blade with a black handle, sharp and clean. The cutting board is thick end-grain maple with a damp towel beneath it so it does not slide. A small white bowl for the diced onion sits at the top of the board. Min-jun holds the onion flat side down, fingertips curled under and knuckles guiding the blade. His knife rocks forward in smooth even strokes, and the dice comes out in neat, uniform quarter-inch cubes. In the top-down shots, his hands stay in the same positions relative to the board between cuts. His delivery is calm, precise, and encouraging, like a good cooking teacher. His white chef jacket has black buttons and a folded sleeve. The backsplash is white subway tile, and a copper pan hangs on a rail behind him in the medium shots. Onion layers glisten and separate realistically. DIALOGUE AND LIP-SYNC: Min-jun speaks in shots one and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Knife tapping on wood, onion crunch. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Clear tutorial. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 bike repair how-to in a shop in Minneapolis, Minnesota, where a mechanic swaps an inner tube. THE SUBJECT: Erik, a 35-year-old white American mechanic with a red beard, a black apron, and grease on his hands. SETTING: A bike repair stand, tools on a pegboard. LIGHTING: Shop fluorescents and window light. CAMERA: Close-ups and medium. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Close-up: tire levers pop the tire bead. 2.5s to 5.0s: Medium: he says: "New tube, a little air." 5.0s to 7.5s: Close-up: the tube goes in. 7.5s to 10.0s: Medium: he spins the wheel and says: "Done." DETAILS: The bike is a steel commuter frame painted dark green, clamped in a blue repair stand by the seat post. The rear wheel is off and held in Erik's hands. The tire is black with tan sidewalls. Two yellow plastic tire levers pry the bead over the rim. A floor pump with a round gauge stands beside him. The pegboard behind holds wrenches, a chain whip, and spare tubes hanging in cardboard sleeves with no readable brand names. Erik's hands have grease on the knuckles and a small bandage on one finger. He works with calm expertise and talks to the viewer like a friendly neighborhood mechanic. In the final shot he spins the wheel and it runs true, with the freewheel ticking. Sunlight falls through a dusty window across the workbench, and a shop cat sleeps on a stool in the background of the medium shots. DIALOGUE AND LIP-SYNC: Erik speaks in shots two and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Tools clinking, pump hissing, freewheel ticking. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Practical how-to. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 plant care tip in a sunny apartment in Miami, Florida. THE SUBJECT: Isabella, a 27-year-old Cuban American woman with curly hair and a green linen shirt. SETTING: Apartment full of plants. LIGHTING: Bright sunlight. CAMERA: Medium and close-ups. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: she says: "Finger test first." 2.5s to 5.0s: Close-up: finger in soil. 5.0s to 7.5s: Close-up: watering can pours. 7.5s to 10.0s: Medium: she says: "Dry means water." DETAILS: The hero plant is a large monstera deliciosa with split leaves in a terracotta pot on a wooden stand by the window. Around it: a fiddle leaf fig, a snake plant, hanging pothos, and a few small succulents on the sill. Isabella's curly dark hair is loose and full, and she wears a thin gold chain. Her green linen shirt has the sleeves rolled up. In shot two, her index finger pushes into dark soil up to the second knuckle and comes out with a little dry dirt on it. In shot three, a matte white metal watering can with a long spout pours a steady stream evenly around the base, soil darkening as it soaks. Water does not overflow. Her delivery is warm and friendly, with a small laugh on the last line. Sunlight comes from the left through sheer white curtains, and leaf shadows fall across the wall behind her. DIALOGUE AND LIP-SYNC: Isabella speaks in shots one and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Water pouring, birds outside. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Friendly tip. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 tech unboxing at a desk in San Francisco, California. THE SUBJECT: Ravi, a 29-year-old Indian American man with glasses and a black hoodie. SETTING: Desk with monitor and LED strip. LIGHTING: Cool LED and daylight. CAMERA: Top-down and medium. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Top-down: he lifts the box lid. 2.5s to 5.0s: Close-up: matte grey headphones. 5.0s to 7.5s: Medium: he puts them on. 7.5s to 10.0s: Medium: he says: "Wow, total silence." DETAILS: The box is a plain matte black rigid box with no readable logos. Inside, the headphones sit in a molded grey tray with a small fabric carry case beside them. The headphones are over-ear, matte grey with soft memory foam cushions and a padded headband, and they stay the same color and shape in every shot. The desk has a large monitor showing a blurred abstract wallpaper, a mechanical keyboard, a small plant, and a purple LED strip glowing behind the monitor. Ravi's glasses are thin black frames and his hoodie is plain black. In shot three he slips the headphones on over his ears, adjusts the headband, and closes his eyes. In shot four his eyebrows lift in genuine surprise and he speaks softly, as if the world just went quiet. The top-down opening shot shows his hands lifting the lid slowly, the tray revealing itself with a slight suction pull. DIALOGUE AND LIP-SYNC: Ravi speaks in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Box sliding, soft room tone. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Clean tech review. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 16:9 science class demo in a high school in Phoenix, Arizona. THE SUBJECT: Mr. Johnson, a 45-year-old Black American teacher with glasses and a sweater vest. SETTING: A classroom lab with students. LIGHTING: Classroom daylight. CAMERA: Wide and close-ups. FRAMING (16:9 widescreen): compose for a landscape screen. Use the full width: place subjects on the left or right third and let the environment breathe on the other side. Keep horizons level and straight. Wide shots should show real geography so the viewer understands where everyone stands, and closer shots should keep the eye line consistent with the wide. No black bars, no split screens. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide: he holds up an egg. 2.5s to 5.0s: Close-up: egg sinks in fresh water. 5.0s to 7.5s: Close-up: egg floats in salt water. 7.5s to 10.0s: Medium: he says: "Salt changes density." DETAILS: The lab has black resin tables, stools, a periodic table poster with no readable text in focus, and a whiteboard with a hand-drawn diagram of two glasses and arrows but no words. Two tall clear glasses of water stand on the front desk: the left one fresh water, the right one salt water with a spoon beside a small bowl of salt. The egg is a white chicken egg, smooth and matte. In shot two it sinks slowly to the bottom of the left glass and rests there. In shot three, placed in the right glass, it floats halfway, bobbing gently. Mr. Johnson's sweater vest is dark green over a light blue shirt with a patterned tie. Students in the wide shot are a diverse group of teenagers leaning forward on their stools, two of them raising eyebrows. His delivery is enthusiastic and clear, like a teacher who loves this demo. DIALOGUE AND LIP-SYNC: He speaks in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Students murmuring, water splash. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Educational. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 1:1 florist clip in a flower shop in Savannah, Georgia. THE SUBJECT: Mae, a 60-year-old white American florist with grey curls and an apron. SETTING: Flower shop counter. LIGHTING: Soft daylight. CAMERA: Top-down and medium. FRAMING (1:1 square): compose for a square feed post. Center the subject with balanced negative space on all four sides, and keep hands, faces, and the key object inside the central 80 percent so nothing important is cut by a rounded corner crop. Favor symmetrical, graphic compositions and top-down or straight-on angles. No borders, no frames inside the frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Top-down: peonies arranged. 2.5s to 5.0s: Top-down: kraft paper wrap. 5.0s to 7.5s: Close-up: twine tied. 7.5s to 10.0s: Medium: she says: "For your mama." DETAILS: The bouquet is soft pink peonies, white garden roses, and sprigs of eucalyptus, arranged in a loose round dome. Mae wraps it in brown kraft paper with a layer of white tissue, folding the edges into neat pleats. The twine is natural jute tied in a simple bow. The shop counter is weathered white wood with buckets of tulips, ranunculus, and hydrangeas behind her. Mae's apron is faded blue denim with pruning shears in the pocket. Her grey curls are pinned back loosely. In the final medium shot she hands the bouquet across the counter to a young man in a plaid shirt whose hands enter frame from the right. Her delivery is warm, with a Southern accent and a knowing smile. Daylight spills through the front window, making the peonies glow. A small bell hangs on the door. DIALOGUE AND LIP-SYNC: Mae speaks in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Paper rustle, scissors snip. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Charming shop moment. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 personal finance tip video filmed in a home office in Charlotte, North Carolina, where a creator explains the 24-hour rule for purchases. THE SUBJECT: Alicia, a 33-year-old Black American creator with a sleek middle-part bob, a mustard sweater, and gold hoops. SETTING: A home office with a plant, a bookshelf, and a laptop. LIGHTING: Soft window light with a warm lamp. CAMERA: Talking head medium with inserts. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: Alicia says: "Want it? Wait a day." 2.5s to 5.0s: Insert: she closes a laptop with a full online cart. 5.0s to 7.5s: Medium: she says: "Most urges fade overnight." 7.5s to 10.0s: Close-up: she smiles and taps a jar of coins. DETAILS: The home office has a white desk, a pothos plant trailing off the shelf, a stack of colorful books, a ceramic mug of pens, and a clear glass jar half full of coins with a handwritten paper label on it. The laptop screen in the insert shows a generic online shopping cart page with product thumbnails and no readable brand names or text. Alicia sits centered in shots one and three, framed from mid-chest up, looking straight into the lens with confident, friendly energy. She uses one small hand gesture on each line: a raised finger on the first line and an open palm on the second. Her sweater, bob, and hoops stay identical. In shot four she leans in, taps the jar lightly, and the coins shift. The pacing feels like a crisp educational short: calm, clear, and trustworthy, with a warm smile. DIALOGUE AND LIP-SYNC: Alicia speaks in shots one and three. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Quiet room, laptop click, coins clinking. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Clean explainer. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 cooking tutorial in a home kitchen in Seattle, Washington, where a chef shows the claw grip while dicing an onion. THE SUBJECT: Chef Min-jun, a 41-year-old Korean American man with short black hair and a white chef jacket, holding a chef knife. SETTING: A wooden cutting board, a yellow onion, and a tiled backsplash. LIGHTING: Bright soft overhead light. CAMERA: Top-down and medium. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: he says: "Tuck your fingertips in." 2.5s to 5.0s: Top-down: claw grip slicing the onion. 5.0s to 7.5s: Top-down: quick even dice. 7.5s to 10.0s: Medium: he says: "Safe and fast." DETAILS: The chef knife is an eight-inch steel blade with a black handle, sharp and clean. The cutting board is thick end-grain maple with a damp towel beneath it so it does not slide. A small white bowl for the diced onion sits at the top of the board. Min-jun holds the onion flat side down, fingertips curled under and knuckles guiding the blade. His knife rocks forward in smooth even strokes, and the dice comes out in neat, uniform quarter-inch cubes. In the top-down shots, his hands stay in the same positions relative to the board between cuts. His delivery is calm, precise, and encouraging, like a good cooking teacher. His white chef jacket has black buttons and a folded sleeve. The backsplash is white subway tile, and a copper pan hangs on a rail behind him in the medium shots. Onion layers glisten and separate realistically. DIALOGUE AND LIP-SYNC: Min-jun speaks in shots one and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Knife tapping on wood, onion crunch. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Clear tutorial. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 16:9 science class demo in a high school in Phoenix, Arizona. THE SUBJECT: Mr. Johnson, a 45-year-old Black American teacher with glasses and a sweater vest. SETTING: A classroom lab with students. LIGHTING: Classroom daylight. CAMERA: Wide and close-ups. FRAMING (16:9 widescreen): compose for a landscape screen. Use the full width: place subjects on the left or right third and let the environment breathe on the other side. Keep horizons level and straight. Wide shots should show real geography so the viewer understands where everyone stands, and closer shots should keep the eye line consistent with the wide. No black bars, no split screens. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide: he holds up an egg. 2.5s to 5.0s: Close-up: egg sinks in fresh water. 5.0s to 7.5s: Close-up: egg floats in salt water. 7.5s to 10.0s: Medium: he says: "Salt changes density." DETAILS: The lab has black resin tables, stools, a periodic table poster with no readable text in focus, and a whiteboard with a hand-drawn diagram of two glasses and arrows but no words. Two tall clear glasses of water stand on the front desk: the left one fresh water, the right one salt water with a spoon beside a small bowl of salt. The egg is a white chicken egg, smooth and matte. In shot two it sinks slowly to the bottom of the left glass and rests there. In shot three, placed in the right glass, it floats halfway, bobbing gently. Mr. Johnson's sweater vest is dark green over a light blue shirt with a patterned tie. Students in the wide shot are a diverse group of teenagers leaning forward on their stools, two of them raising eyebrows. His delivery is enthusiastic and clear, like a teacher who loves this demo. DIALOGUE AND LIP-SYNC: He speaks in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Students murmuring, water splash. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Educational. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 bike repair how-to in a shop in Minneapolis, Minnesota, where a mechanic swaps an inner tube. THE SUBJECT: Erik, a 35-year-old white American mechanic with a red beard, a black apron, and grease on his hands. SETTING: A bike repair stand, tools on a pegboard. LIGHTING: Shop fluorescents and window light. CAMERA: Close-ups and medium. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Close-up: tire levers pop the tire bead. 2.5s to 5.0s: Medium: he says: "New tube, a little air." 5.0s to 7.5s: Close-up: the tube goes in. 7.5s to 10.0s: Medium: he spins the wheel and says: "Done." DETAILS: The bike is a steel commuter frame painted dark green, clamped in a blue repair stand by the seat post. The rear wheel is off and held in Erik's hands. The tire is black with tan sidewalls. Two yellow plastic tire levers pry the bead over the rim. A floor pump with a round gauge stands beside him. The pegboard behind holds wrenches, a chain whip, and spare tubes hanging in cardboard sleeves with no readable brand names. Erik's hands have grease on the knuckles and a small bandage on one finger. He works with calm expertise and talks to the viewer like a friendly neighborhood mechanic. In the final shot he spins the wheel and it runs true, with the freewheel ticking. Sunlight falls through a dusty window across the workbench, and a shop cat sleeps on a stool in the background of the medium shots. DIALOGUE AND LIP-SYNC: Erik speaks in shots two and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Tools clinking, pump hissing, freewheel ticking. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Practical how-to. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 1:1 florist clip in a flower shop in Savannah, Georgia. THE SUBJECT: Mae, a 60-year-old white American florist with grey curls and an apron. SETTING: Flower shop counter. LIGHTING: Soft daylight. CAMERA: Top-down and medium. FRAMING (1:1 square): compose for a square feed post. Center the subject with balanced negative space on all four sides, and keep hands, faces, and the key object inside the central 80 percent so nothing important is cut by a rounded corner crop. Favor symmetrical, graphic compositions and top-down or straight-on angles. No borders, no frames inside the frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Top-down: peonies arranged. 2.5s to 5.0s: Top-down: kraft paper wrap. 5.0s to 7.5s: Close-up: twine tied. 7.5s to 10.0s: Medium: she says: "For your mama." DETAILS: The bouquet is soft pink peonies, white garden roses, and sprigs of eucalyptus, arranged in a loose round dome. Mae wraps it in brown kraft paper with a layer of white tissue, folding the edges into neat pleats. The twine is natural jute tied in a simple bow. The shop counter is weathered white wood with buckets of tulips, ranunculus, and hydrangeas behind her. Mae's apron is faded blue denim with pruning shears in the pocket. Her grey curls are pinned back loosely. In the final medium shot she hands the bouquet across the counter to a young man in a plaid shirt whose hands enter frame from the right. Her delivery is warm, with a Southern accent and a knowing smile. Daylight spills through the front window, making the peonies glow. A small bell hangs on the door. DIALOGUE AND LIP-SYNC: Mae speaks in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Paper rustle, scissors snip. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Charming shop moment. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 plant care tip in a sunny apartment in Miami, Florida. THE SUBJECT: Isabella, a 27-year-old Cuban American woman with curly hair and a green linen shirt. SETTING: Apartment full of plants. LIGHTING: Bright sunlight. CAMERA: Medium and close-ups. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: she says: "Finger test first." 2.5s to 5.0s: Close-up: finger in soil. 5.0s to 7.5s: Close-up: watering can pours. 7.5s to 10.0s: Medium: she says: "Dry means water." DETAILS: The hero plant is a large monstera deliciosa with split leaves in a terracotta pot on a wooden stand by the window. Around it: a fiddle leaf fig, a snake plant, hanging pothos, and a few small succulents on the sill. Isabella's curly dark hair is loose and full, and she wears a thin gold chain. Her green linen shirt has the sleeves rolled up. In shot two, her index finger pushes into dark soil up to the second knuckle and comes out with a little dry dirt on it. In shot three, a matte white metal watering can with a long spout pours a steady stream evenly around the base, soil darkening as it soaks. Water does not overflow. Her delivery is warm and friendly, with a small laugh on the last line. Sunlight comes from the left through sheer white curtains, and leaf shadows fall across the wall behind her. DIALOGUE AND LIP-SYNC: Isabella speaks in shots one and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Water pouring, birds outside. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Friendly tip. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Prompt
A 10-second 9:16 tech unboxing at a desk in San Francisco, California. THE SUBJECT: Ravi, a 29-year-old Indian American man with glasses and a black hoodie. SETTING: Desk with monitor and LED strip. LIGHTING: Cool LED and daylight. CAMERA: Top-down and medium. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Top-down: he lifts the box lid. 2.5s to 5.0s: Close-up: matte grey headphones. 5.0s to 7.5s: Medium: he puts them on. 7.5s to 10.0s: Medium: he says: "Wow, total silence." DETAILS: The box is a plain matte black rigid box with no readable logos. Inside, the headphones sit in a molded grey tray with a small fabric carry case beside them. The headphones are over-ear, matte grey with soft memory foam cushions and a padded headband, and they stay the same color and shape in every shot. The desk has a large monitor showing a blurred abstract wallpaper, a mechanical keyboard, a small plant, and a purple LED strip glowing behind the monitor. Ravi's glasses are thin black frames and his hoodie is plain black. In shot three he slips the headphones on over his ears, adjusts the headband, and closes his eyes. In shot four his eyebrows lift in genuine surprise and he speaks softly, as if the world just went quiet. The top-down opening shot shows his hands lifting the lid slowly, the tray revealing itself with a slight suction pull. DIALOGUE AND LIP-SYNC: Ravi speaks in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Box sliding, soft room tone. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Clean tech review. Keep motion clean and purposeful, with steady exposure and simple, confident blocking in every shot. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. PACING: Each shot is exactly 2.5 seconds. Start each shot mid-action rather than from a frozen pose, and let each spoken line finish fully before its cut. The final shot settles and holds its last half second so the clip ends cleanly. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.
Why creators choose MiniMax H3 Max Turbo
Max quality, faster
H3 Max Turbo is a distilled version of H3 Max tuned to keep its prompt adherence and visual character at higher throughput. It is built for quick creative iteration.
768p at a lower price
Turbo renders the same 480p and 768p sizes as H3 Max at roughly half the standard per-second price, so 768p finals cost less.
Native sound in the Playground
Turbo generates a synced audio track with dialogue and effects described in your prompt. In Fliki's Playground the clip keeps it. In video workflows Fliki lays its own voiceover and music instead.
Multi-shot scenes
Declared hard cuts render cleanly inside one clip, so a 10 second prompt can hold three or four shots with the same cast.
First and last frame control
Pin an opening image, a closing image, or both. Turbo does not take reference images, video, or audio, so frames are the way to lock a look. For references, use H3 Max or H3 Fast.
Long, detailed prompts
Write up to 7,000 characters. Turbo follows full shot lists with lighting, lens, blocking, and sound direction.
5 to 15 second clips
Choose any length from 5 to 15 seconds in one-second steps.
Five native aspect ratios
Render 16:9, 9:16, 1:1, 4:3, or 3:4, each composed natively.
How it works
How to generate a video with MiniMax H3 Max Turbo
Getting a finished clip out of MiniMax H3 Max Turbo takes a few steps inside Fliki. Follow these six.

Write your prompt
Open Fliki and describe the scene. Include the subject, setting, lighting, camera, any dialogue in quotes, and the sounds you want. MiniMax H3 Max Turbo reads up to 7,000 characters, so list shots with timecodes for multi-shot clips.

Select MiniMax H3 Max Turbo as your model
Open the model selector and choose MiniMax H3 Max Turbo. Fliki sends your prompt to MiniMax's model with no extra setup.

Pick your aspect ratio
Choose 16:9 for YouTube, 9:16 for TikTok, Reels, and Shorts, 1:1 for feeds, or 4:3 and 3:4. MiniMax H3 Max Turbo composes each ratio natively.

Set the duration
Pick any length from 5 to 15 seconds. Give multi-shot sequences 10 seconds or more so each shot has room.

Add first or last frame (optional)
Upload an opening frame, a closing frame, or both to lock the look and the motion path. H3 Max Turbo uses frame images rather than reference images.

Select resolution and generate
Pick 768p for final output or 480p for the cheapest take, then hit Generate. Need references or 2K? Try H3 Max or MiniMax H3.
AI MODEL GALLERY
Built on the best AI models - ready inside Fliki
Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.
















MiniMax H3 Max Turbo FAQ
Frequently asked questions
Everything you need to know about generating with MiniMax H3 Max Turbo inside Fliki.
MiniMax H3 Max Turbo is a distilled, speed-focused variant of H3 Max. It generates 5 to 15 second clips at 480p or 768p from text or an opening image, with optional end-frame guidance and synchronized audio. It became available through Runware on September 2, 2026.
No. H3 Max Turbo is available on paid Fliki plans starting with Basic.
Fliki offers 768p (1344x768 in 16:9) and 480p (832x480 in 16:9). For 2K, use MiniMax H3.
Yes. Turbo renders a synced audio track, and clips made in the Playground keep it. When Turbo renders scenes inside a Fliki video workflow, its track is removed and Fliki adds your chosen voiceover, music, and sound effects instead.
Turbo accepts a first frame, a last frame, or both. It does not accept reference images, video, or audio. For those, use H3 Max, MiniMax H3, or H3 Fast.
From 5 to 15 seconds in one-second steps, at 24 frames per second.
Pick MiniMax H3 when you need the highest resolution, since it is the only tier with a 2K (1440p) option on Fliki. Pick H3 Max for the most faithful 768p results with the full reference set. Pick H3 Max Turbo for 768p or 480p drafts at a lower price when a first or last frame is all the guidance you need. Pick H3 Fast for the cheapest 480p previews that still accept image, video, and audio references.
Turbo covers 768p at a low per-second price and follows long prompts well, which suits workflows that render many scenes per video. You can pick another model in those workflows at any time.
Still curious?
Try Fliki free in your browser, no credit card required.
Start free→More from Fliki
AI video models
Discover more
Tools
Discover features
Generate your next video with MiniMax H3 Max Turbo.
Fast 768p renders with native sound and first and last frame control. Free to start, no credit card required.
Generate your first video freeFree forever plan · No credit card required · Cancel anytime

