video model · by MiniMax

MiniMax H3 Fast AI Video Generator

Preview ideas quickly with MiniMax H3 Fast, the 480p speed tier of the Hailuo 3.0 family. It keeps the full H3 reference set of 9 images, 3 videos, and 3 audio clips plus native sound, at the lowest price of any H3 tier on Fliki. Compare it side by side with all our AI video models before you render.

Generated with MiniMax H3 Fast

A handful of MiniMax H3 Fast clips generated inside Fliki. No edits, no post.

Prompt

A 10-second 9:16 before and after reveal at a neighborhood dog grooming salon in Austin, Texas, as a groomer finishes a very fluffy goldendoodle. THE SUBJECT: Kayla, a 27-year-old white American groomer with a high blonde ponytail, a teal grooming smock, and a slicker brush in her hand. The dog is Biscuit, a large apricot goldendoodle with a teddy bear face, first shown shaggy and matted, then freshly cut with a round fluffy head and a small blue bandana. SETTING: A bright grooming room with a stainless steel table and grooming arm, a blow dryer on a stand, pegboard walls of combs and scissors, and a big window. LIGHTING: Clean bright daylight plus overhead LED panels, even and cheerful. CAMERA: Phone-style handheld, playful, with quick framing and a close-up of the dog. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium shot: shaggy Biscuit on the table, fur in his eyes. Kayla looks at the camera and says: "Before. A total mop." 2.5s to 5.0s: Close-up: the blow dryer blasts, fur flies up in a big fluffy cloud, Biscuit squints happily. 5.0s to 7.5s: Close-up: Kayla ties the blue bandana around Biscuit's neck, his fresh round face now clean. 7.5s to 10.0s: Medium shot: she steps aside with jazz hands. She says: "After. Handsome boy!" Biscuit wags. DETAILS: Biscuit sits patiently on the table the whole time, tongue out, tail thumping. The shaggy fur in shot one is visibly longer and tangled, and in shot four it is short, even, and fluffed, with a clean round muzzle and neatly rounded ears. The blue bandana has a small white paw print pattern. Kayla has a pair of curved grooming shears tucked in her smock pocket, handles visible, and a few apricot hairs stuck to her sleeves. The background pegboard is colorful with pink, green, and purple brushes hanging in neat rows, and a small sign on the wall shows a paw print drawing without readable words. Kayla's energy is big and funny, as if showing off her best work to followers. DIALOGUE AND LIP-SYNC: Kayla speaks in shots one and four with a bright Texas accent, excited and funny. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Small tiled room. The loud whoosh of the blow dryer. Tags jingling. One happy bark at the end. Nails tapping on the steel table. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Fun short-form social clip with a clear transformation. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 hype clip at a strength gym in Atlanta, Georgia, as a coach cheers a lifter through a personal record deadlift. THE SUBJECT: Denise, a 42-year-old Black American strength coach with short locs tied back, a black tank top, and a stopwatch around her neck. The lifter Sam, a 30-year-old Vietnamese American woman with a high bun, chalked hands, a maroon sports bra and black shorts, and a wide leather lifting belt. The hero prop is a loaded barbell with red and blue bumper plates. SETTING: A warehouse gym with rubber flooring, chalk buckets, squat racks, and a big roll-up door open to daylight. LIGHTING: Mixed daylight from the door and overhead industrial lights, bold contrast, chalk dust visible in the air. CAMERA: Low angle handheld for power, one close-up of hands, fast energy. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Low angle: Sam grips the bar, Denise crouches beside her and says: "Big breath. Push the floor." 2.5s to 5.0s: Close-up of chalky hands tightening on the knurled bar, chalk puffing up. 5.0s to 7.5s: Low angle wide: Sam drives the bar up to lockout, face straining, plates bending the bar slightly. 7.5s to 10.0s: Medium: she drops the bar, it bounces on the rubber. Sam yells: "That is a PR!" Denise hugs her. DETAILS: The bar is loaded with two red plates and one blue plate per side, and the load stays the same in every shot. Sam wears black knee-high lifting socks and flat black shoes. Denise holds a small white towel over her shoulder. In the background, two other lifters pause their sets to watch, one clapping in shot four. A whiteboard on the wall is covered in hand-drawn tally marks and arrows but no readable text. Chalk dust hangs in the shafts of daylight from the roll-up door, drifting slowly. Sam's face shows real strain in shot three: clenched jaw, flushed cheeks, a vein at the temple, eyes fixed forward. DIALOGUE AND LIP-SYNC: Denise speaks low and focused in shot one. Sam shouts with joy in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Echoing warehouse ambience, clanking plates, a loud rubber thud of the dropped barbell, other lifters cheering in the background. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Energetic fitness social content. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 close and cozy coffee shop clip in Portland, Oregon, where a barista pours a rosetta and hands it across the counter. THE SUBJECT: Jonah, a 26-year-old white American barista with a curly red beard, a mustard beanie, a black apron over a striped long-sleeve shirt, and a silver steaming pitcher. The customer Aaliyah, a 23-year-old Somali American woman in a lavender hijab and a cream knit sweater. SETTING: A small cafe with a matte black espresso machine, a wooden counter, a chalkboard menu, and rain on the front window. LIGHTING: Soft grey rainy daylight from the window with warm pendant lights above the counter. CAMERA: Top-down and close handheld shots, steady on the pour. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Close-up: steam wand hisses into milk in the pitcher, Jonah taps it on the counter. 2.5s to 5.0s: Top-down: he pours into a white cup of espresso, wiggling the pitcher, and a rosetta leaf forms. 5.0s to 7.5s: Medium: he slides the cup across. He says: "Oat milk, extra hot." 7.5s to 10.0s: Close-up on Aaliyah smiling at the leaf. She says: "Too pretty to drink." DETAILS: The espresso is a rich crema-topped double in a wide white ceramic cup on a white saucer, with a small silver spoon on the saucer. The rosetta has seven clear leaves and a pulled-through stem, and it stays intact when the cup slides across the counter. Jonah's movements are practiced and smooth. A glass jar of biscotti and a small succulent sit on the counter. Aaliyah holds her phone in one hand as if about to take a picture of the cup. Rain streaks down the window behind her in slow wandering lines, and a blurred red umbrella passes outside. DIALOGUE AND LIP-SYNC: Jonah speaks casually in shot three, Aaliyah warmly in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Loud steam wand hiss then the tap of the pitcher. Espresso grinder in the background. Rain on glass. Soft cafe chatter. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Cozy, satisfying short-form coffee content. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 move-in day clip at a college dorm in Madison, Wisconsin, as a mom helps her son carry the last box and says goodbye. THE SUBJECT: Linda, a 49-year-old Filipino American mom with a shoulder-length bob, a red university sweatshirt, and sunglasses pushed up on her head. Her son Nate, 18, with messy black hair, a grey T-shirt, and basketball shorts, carrying a cardboard box labeled "DESK STUFF" in black marker. SETTING: A dorm hallway with cinder block walls, name tags on doors, and a small room with a lofted bed and a mini fridge. LIGHTING: Fluorescent hallway light plus warm late-afternoon sun through the room's window. CAMERA: Handheld, family home-video energy, close and warm. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: Nate backs through the door with the box, Linda holding the door open. 2.5s to 5.0s: Close-up: the box lands on the desk. Linda says: "That is the last one." 5.0s to 7.5s: Medium two-shot: she hugs him tight, he rolls his eyes but hugs back. 7.5s to 10.0s: Close-up on Nate at the door as she leaves. He says: "Text you tonight, Mom." DETAILS: The hallway is busy with other families: a dad carrying a lamp, a girl rolling a pink suitcase, and a bulletin board covered in colorful paper cutouts with no readable text. The dorm room has a navy comforter on the lofted bed, a desk lamp, and a string of photo clips on the wall with a few printed family photos. Linda's sunglasses stay on her head in every shot. Nate is taller than his mother and bends down to hug her. The "DESK STUFF" marker label on the box stays the same, facing the camera in shots one and two. DIALOGUE AND LIP-SYNC: Linda speaks in shot two, bright but emotional. Nate answers softly in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Busy hallway with distant voices, a rolling cart, a door thud, the cardboard box landing, a small sniffle during the hug. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Warm, relatable family moment. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 thrift haul try-on in a small apartment in Chicago, Illinois, where a creator shows off a vintage leather jacket find. THE SUBJECT: Mateo, a 25-year-old Colombian American creator with a wavy black mullet, a small silver nose ring, a white ribbed tank top, and baggy light-wash jeans. The hero prop is a worn brown vintage leather bomber jacket with a shearling collar and a paper price tag hanging from the sleeve that reads "$18". SETTING: A bedroom with a full-length mirror, a clothing rack, string lights, and a window with city rooftops. LIGHTING: Soft window daylight with warm string lights in the background. CAMERA: Mirror and front-facing handheld phone framing, quick and playful. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: Mateo holds up the jacket to the camera. He says: "Eighteen dollars. Look at this." 2.5s to 5.0s: Close-up: his hand flips the price tag and runs over the cracked leather and shearling collar. 5.0s to 7.5s: Medium: he swings the jacket on and pops the collar in the mirror. 7.5s to 10.0s: Close-up: he grins at the lens. He says: "Main character energy." DETAILS: The leather jacket is clearly worn in: cracked creases at the elbows, a faded patch on the left shoulder, and a brass zipper with a small leather pull. The shearling collar is cream and slightly matted. The price tag is a small manila paper tag on white string, handwritten in marker, and it stays on the right sleeve cuff in every shot. A clothing rack behind him holds colorful vintage tees, a corduroy blazer, and a denim jacket. Mateo's energy is excited, expressive, a little dramatic in a funny way, with a quick eyebrow raise at the end. DIALOGUE AND LIP-SYNC: Mateo speaks to the lens in shots one and four, excited and playful. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Small room tone, leather creaking, the jacket swishing on, a hanger clicking on the rack, faint city traffic outside. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Casual, fun fashion social clip. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 16:9 lively food truck scene at a night market in Los Angeles, California, as a cook serves birria tacos to a couple. THE SUBJECT: Chef Luis, a 38-year-old Mexican American cook with a black bandana, a thick mustache, and a white T-shirt under a black apron. The customers Emily, a 29-year-old white American woman with a blonde pixie cut and a denim jacket, and her partner Kevin, a 30-year-old Chinese American man with round glasses and a hoodie. The hero food is three birria tacos with a cup of red consomme. SETTING: A food truck with a glowing service window, string lights, folding tables, and a crowd at a night market in a parking lot. LIGHTING: Warm truck window light and colorful string lights against a dark night. CAMERA: Wide establishing, then handheld close shots at the window. FRAMING (16:9 widescreen): compose for a landscape screen. Use the full width: place subjects on the left or right third and let the environment breathe on the other side. Keep horizons level and straight. Wide shots should show real geography so the viewer understands where everyone stands, and closer shots should keep the eye line consistent with the wide. No black bars, no split screens. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide: the glowing truck on the left third, a line of people, string lights overhead. 2.5s to 5.0s: Close-up: Luis dips a tortilla in red consomme and flips it on the griddle, sizzling. 5.0s to 7.5s: Medium at the window: he hands out the tray. He says: "Dip it, trust me." 7.5s to 10.0s: Medium: Emily dips and bites, eyes widen. She says: "Oh, wow." Kevin laughs. DETAILS: The food truck is painted matte black with a hand-painted orange and pink flower pattern around the service window, and no readable lettering. A menu board inside the window shows only photos. Emily and Kevin stand on the right third of the frame at the window, sharing one paper tray. The consomme is deep red with a shine of fat on top, garnished with chopped onion and cilantro. The tacos have crispy red-stained edges and melted cheese pulling as Emily bites. The crowd behind them moves naturally, and a kid with a glowing balloon walks past in the wide shot. DIALOGUE AND LIP-SYNC: Luis speaks with warm confidence in shot three. Emily reacts in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Busy outdoor crowd, griddle sizzle, a generator hum, distant cumbia from another stall, the crunch of a taco. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Vibrant street food social content. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 1:1 family science experiment at a kitchen table in Denver, Colorado, where a dad and his daughter make a baking soda volcano erupt. THE SUBJECT: Andre, a 40-year-old Black American dad with a short beard, a grey henley, and safety goggles on his forehead. His daughter Zoe, 8, with two puffs, a yellow T-shirt, and oversized safety goggles. The hero prop is a small clay volcano painted brown on a baking tray. SETTING: A bright kitchen with a white table covered in newspaper, a box of baking soda, a bottle of vinegar, and red food coloring. LIGHTING: Bright even daylight from a window behind the camera. CAMERA: Top-down and front shots, steady, centered. FRAMING (1:1 square): compose for a square feed post. Center the subject with balanced negative space on all four sides, and keep hands, faces, and the key object inside the central 80 percent so nothing important is cut by a rounded corner crop. Favor symmetrical, graphic compositions and top-down or straight-on angles. No borders, no frames inside the frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Front shot: Zoe holds up the vinegar. She says: "Ready, Dad?" 2.5s to 5.0s: Top-down: Andre adds red coloring to the vinegar, Zoe pours it into the crater. 5.0s to 7.5s: Top-down: red foam bubbles up and spills down the volcano sides. 7.5s to 10.0s: Front shot: both laugh with hands up. Andre says: "We did science!" DETAILS: The volcano is lumpy and hand-made, about a foot tall, painted brown with a few green painted trees on its base. Zoe's goggles are a little too big and slide down her nose once. Andre wears his goggles on his forehead until shot two, where he pulls them down. The baking soda box and vinegar bottle stay on the left side of the tray in each overhead shot. The red foam is thick and bubbly, flowing down in slow rivers and pooling on the newspaper. A fridge with crayon drawings held by magnets is visible in the background of the front shots. DIALOGUE AND LIP-SYNC: Zoe speaks in shot one, excited. Andre cheers in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Kitchen room tone, a loud fizzing hiss as the foam rises, paper rustling, laughter. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Fun, bright family educational clip. Keep the composition bold and readable, with clear shapes and strong subject separation, so it reads well at small sizes. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose MiniMax H3 Fast

Built for quick iteration

H3 Fast is the speed tier of the H3 family. Use it to test a concept, a camera move, or a line of dialogue before you spend credits on a higher-resolution render.

Lowest H3 price on Fliki

At 480p, H3 Fast costs less per second than any other H3 tier on Fliki, which makes batch drafts and A/B variations cheap enough to run freely.

Full reference support

Unlike H3 Max Turbo, H3 Fast keeps multi-modal references: up to 9 images, 3 videos, and 3 audio clips. Lock a character, copy a motion, or keep a voice even in drafts.

Native audio included

Dialogue, effects, and ambience come back in the same generation. Hear how a line lands before you commit to a final render.

Multi-shot in a single clip

Declared hard cuts render cleanly inside one H3 Fast generation, so you can block out a whole sequence and check its rhythm at draft speed.

4 to 15 second clips

H3 Fast is the only H3 tier that goes down to 4 seconds. Short loops and quick reactions cost less, and the ceiling still reaches 15 seconds.

First and last frame control

Pin an opening frame, a closing frame, or both. It is a quick way to test transitions between scenes in a Fliki project.

Long prompts, five ratios

Write up to 7,000 characters and render in 16:9, 9:16, 1:1, 4:3, or 3:4, each composed natively at 480p.

How it works

How to generate a video with MiniMax H3 Fast

Getting a finished clip out of MiniMax H3 Fast takes a few steps inside Fliki. Follow these six.

Fliki prompt input with a multi-shot description for the MiniMax H3 Fast AI video generator
Step 1

Write your prompt

Open Fliki and describe the scene. Include the subject, setting, lighting, camera, any dialogue in quotes, and the sounds you want. MiniMax H3 Fast reads up to 7,000 characters, so list shots with timecodes for multi-shot clips.

Fliki model selector with MiniMax H3 Fast chosen for AI video generation
Step 2

Select MiniMax H3 Fast as your model

Open the model selector and choose MiniMax H3 Fast. Fliki sends your prompt to MiniMax's model with no extra setup.

Choose an aspect ratio for MiniMax H3 Fast video generation on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube, 9:16 for TikTok, Reels, and Shorts, 1:1 for feeds, or 4:3 and 3:4. MiniMax H3 Fast composes each ratio natively.

Set video duration on the Fliki slider for MiniMax H3 Fast
Step 4

Set the duration

Pick any length from 4 to 15 seconds. Keep drafts short, and give multi-shot sequences 10 seconds or more.

Upload optional guidance media for MiniMax H3 Fast on Fliki
Step 5

Add references (optional)

Upload reference images, video, or audio to lock a character, a camera move, or a voice, or pin a first and last frame instead. The two approaches are used separately.

Pick output resolution and generate video with MiniMax H3 Fast on Fliki
Step 6

Generate

H3 Fast renders at 480p only, so there is nothing to pick. Hit Generate and move the winning take to H3, H3 Max, or H3 Max Turbo if you need more resolution.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

MiniMax H3 Fast FAQ

Frequently asked questions

Everything you need to know about generating with MiniMax H3 Fast inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
MiniMax H3 Fast · Free forever plan

Draft your next video with MiniMax H3 Fast.

Quick 480p clips with native audio and full reference support at the lowest H3 price. Free to start, no credit card required.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime