video model · by MiniMax

MiniMax H3 AI Video Generator

Create 2K videos with native stereo sound using MiniMax H3, the Hailuo 3.0 model. Direct multi-shot scenes with dialogue, guide them with up to 9 reference images, 3 videos, and 2 audio clips, and render up to 15 seconds, all inside Fliki. Compare it side by side with all our AI video models before you render.

Generated with MiniMax H3

A handful of MiniMax H3 clips generated inside Fliki. No edits, no post.

Prompt

A 10-second 9:16 intimate craft film inside a small independent watch workshop in Brooklyn, New York, showing an old master and his apprentice finishing a mechanical movement by hand. THE SUBJECT: Walter, a 68-year-old white American watchmaker with a trimmed silver beard, wire-rimmed half glasses, a navy cardigan over a pale blue oxford shirt with the sleeves rolled twice, and a steel loupe clipped to the right lens of his glasses. His apprentice Imani, a 24-year-old Black American woman with short natural coils, small gold hoop earrings, and a charcoal canvas work apron over a cream turtleneck. The hero prop is a single open watch movement about the size of a quarter, brass plates with blued steel screws and a ruby jewel bearing, resting on a small grey rubber cushion. SETTING: A narrow second-floor workshop with a scarred walnut bench, a green leather bench pad, rows of tiny labeled drawers, a brass bench lamp, tweezers and screwdrivers standing in a wooden block, and a tall window with old wavy glass looking onto a brick wall across the street. Dust motes hang in the air. LIGHTING: Soft late-afternoon daylight through the window from camera left, mixed with the warm pool of the bench lamp directly over the movement. Skin tones stay natural and warm, the brass glows, and the blued screws catch a cool blue highlight. Deep but not crushed shadows in the room behind. CAMERA: A mix of locked macro close-ups and a gentle handheld medium. Shallow depth of field on the macro shots with a slow focus pull from tweezer tip to jewel. 24fps with natural motion blur. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Extreme macro close-up of the movement. Walter's steady fingers, steel tweezers in the right hand, lower a tiny blued screw into place. The balance wheel begins to oscillate, catching light each swing. 2.5s to 5.0s: Medium two-shot across the bench. Walter leans back and lifts the loupe; Imani leans in on his left. He says to her: "Listen. It is breathing now." 5.0s to 7.5s: Close-up of Imani, face lit by the lamp, the loupe now in her eye, a slow smile. She says quietly: "Five beats a second." 7.5s to 10.0s: Macro insert of the finished movement from above, balance wheel ticking steadily, then Walter's hand slides a polished steel case back over it and presses it shut with a small click. DETAILS: Walter works with the patience of fifty years: his hands never tremble, and he breathes out slowly before each placement. Imani sits on a lower stool to his left, forearms on the bench edge, watching his tweezers rather than his face until he speaks to her. The brass lamp head is angled low, so the movement sits in a bright circle while the drawers behind fall into soft shadow. DIALOGUE AND LIP-SYNC: Two lines only. Walter speaks in shot two in a low, gravelly, unhurried New York voice. Imani answers in shot three, softly, almost to herself, with warmth. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: A fast, delicate mechanical ticking that starts when the balance wheel moves in shot one and continues under everything. The soft metallic tap of tweezers on brass. The creak of a wooden stool. Distant muffled city traffic through the old window. The final crisp click of the case back closing. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Quiet, respectful documentary portrait of craft, like a short film for a heritage brand. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 creator-style skincare ad filmed in a bright bathroom in Scottsdale, Arizona, where a woman demonstrates a vitamin C serum as part of her morning routine. THE SUBJECT: Priya, a 31-year-old Indian American woman with long dark wavy hair clipped up in a tortoiseshell claw clip, a few loose strands framing her face, bare skin with visible natural texture and a faint dusting of freckles across the nose, a soft oatmeal ribbed tank top, and thin gold stud earrings. The hero product is a 30 ml amber glass dropper bottle with a white rubber bulb and a matte white label that reads "DAYBREAK C" in simple black type. SETTING: A modern bathroom with a white oak floating vanity, a round mirror with a thin brass frame, pale terracotta zellige tiles, a small trailing pothos on the counter, a folded waffle hand towel, and a frosted window behind her. LIGHTING: Bright, clean morning daylight through the frosted window, diffused and flattering, with a soft reflection in the mirror. The amber glass glows when backlit. No harsh specular hot spots on the skin. CAMERA: Handheld phone-style framing at arm's length for the talking shots, as if she is filming herself, plus two steady product close-ups. Natural slight sway, fast autofocus, 24fps. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium close-up of Priya facing the camera, holding the amber bottle beside her cheek. She smiles and says: "My three second morning step." 2.5s to 5.0s: Macro close-up: her fingers squeeze the white bulb and three golden drops fall into her palm, each drop catching the window light. 5.0s to 7.5s: Close-up of her face as she presses the serum into her cheeks with both palms, the skin taking on a soft dewy sheen, eyes closed briefly. 7.5s to 10.0s: Medium close-up again, she opens her eyes, taps the bottle, and says: "Glow first, coffee second." DETAILS: Behind her on the counter: a ceramic cup holding a bamboo toothbrush, a small linen-scented candle, and a folded white washcloth. Her claw clip stays in the same position in every shot. The label on the bottle faces the camera whenever she holds it, and the lettering stays sharp and correctly spelled. She wears no makeup, and her skin reads real: small pores on the nose, a faint pink on the cheeks, a few baby hairs at the hairline. DIALOGUE AND LIP-SYNC: Priya speaks directly to the lens in shots one and four, relaxed and friendly, in a natural American accent, like she is talking to a friend on FaceTime. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Small bathroom room tone with a slight tile echo. The soft squeeze of the rubber bulb and the tiny drip into her palm. The gentle pat of hands on skin. The light clink of glass set down on the counter. Faint birdsong outside the window. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Authentic user generated content look with premium product detail, honest skin, not over-retouched. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 street food vignette at a halal cart on a busy corner in Jackson Heights, Queens, at dusk, following a vendor cooking chicken over rice for a regular customer. THE SUBJECT: Tariq, a 45-year-old Egyptian American cart vendor with a close-cropped salt and pepper beard, a black knit beanie, a grey hoodie under a red apron with grease spots, and clear food-safe gloves. His customer Danny, a 27-year-old Puerto Rican American man in a green Carhartt work jacket, a white hard hat tucked under his arm, and paint-flecked boots. The hero food is chicken over yellow rice in a foil tray with shredded lettuce, tomato, and generous white and red sauce. SETTING: A stainless steel food cart with a sizzling flat-top griddle, squeeze bottles of white and red sauce, a stack of foil trays, and a small glowing menu board with photos only. Behind them: the elevated 7 train tracks, a row of shop signs, yellow cabs, and pedestrians blurred in the background. LIGHTING: Blue hour sky mixed with warm tungsten from the cart's interior bulb and the orange glow of shop signs. Steam from the griddle catches the light in thick drifting clouds. Faces lit warm from the cart side and cool from the sky side. CAMERA: Handheld street documentary feel, slightly low, close to the action, with one static wide establishing shot. 24fps, natural motion blur on the spatula moves. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Close-up of the griddle: Tariq chops chicken with two metal spatulas in rapid rhythm, steam and sizzle rising, spices darkening on the edges. 2.5s to 5.0s: Medium shot over the cart counter. Danny steps up, nods. Tariq grins and says: "The usual, extra white sauce?" 5.0s to 7.5s: Insert close-up: Tariq spoons chicken onto the rice and zigzags white sauce, then a thin line of red sauce across the top. 7.5s to 10.0s: Medium two-shot. Tariq hands over the closed foil tray in a plastic bag. Danny says: "You know me, boss." A train rumbles overhead. DETAILS: Background life keeps moving the whole time: a woman in a sari walks past with shopping bags, two teenagers wait at a bus stop, and a delivery cyclist rolls by on the curb side. The cart's steel surfaces are scratched and dented from years of use. Steam drifts left across frame in every shot because a light breeze comes from the right. DIALOGUE AND LIP-SYNC: Tariq speaks warmly with a light Egyptian accent in shot two. Danny answers with a relaxed New York accent in shot four, grinning. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Loud griddle sizzle and the rapid metallic clack of spatulas. Squeeze bottles squelching. City ambience: car horns in the distance, footsteps, fragments of conversation, and in shot four the rising roar and rattle of the 7 train passing on the tracks above. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Warm, human street documentary, like a food segment from a travel series. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 wedding reception moment in a converted barn in Vermont, where a nervous best man gives a short toast and the room reacts. THE SUBJECT: Marcus, a 34-year-old Black American man with a short fade, a neat beard, a charcoal three-piece suit with a sage green tie and a small white rose boutonniere, holding a champagne flute in his right hand and a folded index card in his left. The groom Ben, a 33-year-old white American man with curly auburn hair and a navy suit with the same sage tie, seated at the head table beside his bride Hannah, a 31-year-old Korean American woman with a sleek low bun and a simple silk slip dress. SETTING: A rustic barn reception with exposed wooden beams, strings of warm Edison bulbs overhead, long farm tables with eucalyptus garlands, candles in glass, and about sixty guests at round tables, mostly out of focus. LIGHTING: Warm, low evening interior light from the string bulbs and candles, with golden bokeh in the background and a soft spotlight feel on Marcus from a nearby lantern. Champagne glows amber. CAMERA: Wedding videographer style: a steady medium on Marcus with a slight push-in, reaction close-ups on a longer lens, and a wide of the room. 24fps. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium shot of Marcus standing, lifting the flute, glancing at his card and then folding it away. He says: "I had notes. Forget the notes." 2.5s to 5.0s: Close-up reaction of Ben and Hannah at the head table, both laughing, Hannah squeezing Ben's arm. 5.0s to 7.5s: Close-up on Marcus, eyes getting glassy, voice softer: "He found his person. Finally." 7.5s to 10.0s: Wide shot of the barn as sixty guests rise and lift their glasses; Marcus raises his flute high and the room cheers. DETAILS: Blocking: Marcus stands at the end of the head table, Ben and Hannah seated to his right, so his eye line in shots one and three goes slightly screen right toward them. Guests in the background are a mix of ages, some holding phones low to film. Candle flames flicker gently and stay consistent in position between the wide and the close shots. DIALOGUE AND LIP-SYNC: Only Marcus speaks, in shots one and three, with a warm baritone and a small nervous laugh in the first line. Nobody else says words; guests only laugh and cheer. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: A large wooden room with a gentle echo. A spoon tapping a glass at the very start. Light laughter building in shot two. A quiet pause in shot three with a single sniffle. A big cheer and the clink of many glasses in shot four. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Emotional, candid wedding film, cinematic but real. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 day-in-the-life clip at a small-town fire station in Ohio at dawn, as a firefighter runs the morning equipment check on the engine. THE SUBJECT: Lt. Rosa Delgado, a 39-year-old Mexican American firefighter with a dark low ponytail, a navy department T-shirt tucked into navy work pants with a black belt, and tan bunker pants with suspenders hanging at her hips. Her rookie Tyler, a 22-year-old white American man with a buzz cut, a matching navy T-shirt, and a clipboard. The hero prop is the red engine with polished chrome, gauges, and a coiled yellow hose. SETTING: An open apparatus bay with sealed grey concrete floors, the red engine parked nose-out, turnout gear hanging on open lockers with helmets on top, a bulletin board, and the big roll-up door half open to a quiet street with a bakery sign across the road. LIGHTING: Cool blue dawn light through the half-open bay door mixed with overhead fluorescent tubes, and a warm sunrise sliver beginning to hit the chrome by the final shot. CAMERA: Steady handheld, observational, eye level, with two close inserts. 24fps. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide shot of the bay: Rosa walks along the engine, running a hand over the chrome rail, Tyler following with the clipboard. 2.5s to 5.0s: Close-up insert: Rosa opens a side compartment, taps a pressure gauge with one finger, and the needle settles in the green. 5.0s to 7.5s: Medium two-shot. She turns to Tyler and says: "Check it twice. Every morning." 7.5s to 10.0s: Medium on Tyler making a mark on the clipboard, he nods and says: "Yes, Lieutenant." Sunlight hits the chrome behind him. DETAILS: The engine number "7" is painted in gold on the cab door and stays identical in each shot. Rosa moves with calm authority and a practiced economy of motion. Tyler is eager and a little stiff, standing a half step behind her. A coffee mug steams on a folding table near the lockers, and a department flag hangs on the back wall. DIALOGUE AND LIP-SYNC: Rosa speaks firmly but kindly in shot three. Tyler answers crisply in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Large concrete bay reverb. Boots on concrete. The metal latch and hinge of the compartment. A faint radio chatter from a scanner on a desk, unintelligible. Birds and a single passing car outside. The scratch of a pen on paper. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Grounded, respectful public service documentary. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 16:9 cinematic family moment on a horse ranch in Montana at sunrise, where a father watches his teenage daughter ride out for the first time on her own horse. THE SUBJECT: Cole, a 52-year-old white American rancher with a weathered face, a grey stubble beard, a sweat-stained tan felt cowboy hat, a faded denim jacket with a sherpa collar, and leather work gloves. His daughter June, 16, with a long blonde braid, a burgundy puffer vest over a flannel shirt, jeans, and brown cowboy boots, riding a chestnut quarter horse named Dusty with a white blaze and a worn brown saddle. SETTING: A wooden split-rail corral opening onto a vast golden grassland, frost on the grass, low mist in the valley, and the Rocky Mountain foothills beyond. A weathered red barn on the left edge of frame. LIGHTING: Low golden sunrise from behind the mountains, strong backlight rimming the horse, the hat, and the mist, with breath visible in the cold air. Long shadows across the frost. CAMERA: Wide anamorphic-style framing on a tripod for the landscape, a slow tracking shot alongside the horse, and a tight profile of Cole. 24fps. FRAMING (16:9 widescreen): compose for a landscape screen. Use the full width: place subjects on the left or right third and let the environment breathe on the other side. Keep horizons level and straight. Wide shots should show real geography so the viewer understands where everyone stands, and closer shots should keep the eye line consistent with the wide. No black bars, no split screens. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Extreme wide shot: June rides Dusty through the open corral gate at a walk, Cole standing by the fence post on the left third, mountains on the right. 2.5s to 5.0s: Tracking side shot of June as Dusty breaks into a trot, braid bouncing, breath clouds from horse and rider. 5.0s to 7.5s: Close-up profile of Cole leaning on the fence, a small proud smile. He says quietly: "Easy on her, Dusty." 7.5s to 10.0s: Wide shot from behind Cole: June turns in the saddle and calls back: "I got this, Dad!" then rides on toward the light. DETAILS: Dusty moves with real equine weight: the hindquarters push, the head bobs with each step, the mane lifts and falls, and the tail swishes once. June sits the saddle with confident posture, reins in her left hand. Cole's gloves rest on the top rail, and his hat brim casts a shadow over his eyes except in the close-up, where the low sun catches the grey in his stubble. DIALOGUE AND LIP-SYNC: Cole speaks almost under his breath in shot three, low and gentle. June calls out loudly and happily across the distance in shot four, her voice carrying with open air, no echo. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Wide open air ambience with a light wind through grass. Hooves on frozen dirt then soft thuds on grass, tack jingling, the horse snorting. A meadowlark singing in the distance. The wooden gate creaking at the start. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Epic, warm, Americana cinematic look with rich natural color. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 1:1 overhead-friendly craft clip in a ceramics studio in Asheville, North Carolina, where an instructor helps a student center clay on the wheel. THE SUBJECT: Grace, a 58-year-old Japanese American ceramicist with a chin-length grey bob, clay-smeared forearms, a denim work shirt with rolled sleeves, and a rust linen apron. Her student Leo, a 29-year-old white American man with a curly brown mop, a black T-shirt, and a small tattoo of a fern on his forearm. The hero prop is a lump of wet terracotta clay on a spinning grey wheel head. SETTING: A sunlit studio with shelves of bisque-fired bowls, buckets of slip, sponges, a bucket of water, and a big north-facing window. LIGHTING: Soft even north light from the window, gentle and shadowless, making the wet clay glisten. Warm skin, true terracotta orange. CAMERA: Alternating straight-down overhead shots and eye-level close-ups, steady, on a tripod. 24fps. FRAMING (1:1 square): compose for a square feed post. Center the subject with balanced negative space on all four sides, and keep hands, faces, and the key object inside the central 80 percent so nothing important is cut by a rounded corner crop. Favor symmetrical, graphic compositions and top-down or straight-on angles. No borders, no frames inside the frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Overhead shot of the wheel: Leo's hands wobble on the spinning lump, the clay pushing off center. 2.5s to 5.0s: Eye-level close-up: Grace places her hands over his and says calmly: "Brace your elbows. Let it come." 5.0s to 7.5s: Overhead shot again: under their joined hands the clay rises into a smooth, perfectly centered cone, water shining. 7.5s to 10.0s: Close-up of Leo's face breaking into a grin. He says: "Oh, I felt that." Grace laughs softly. DETAILS: The wheel spins counterclockwise at a steady medium speed in every overhead shot. A shallow bowl of cloudy water sits at the right edge of the splash pan, and a yellow sponge rests beside it. Clay slip streaks their wrists and stays in the same places between shots. Grace's hands are strong and veined, with short unpolished nails; Leo's hands are wider and less sure until the center locks. DIALOGUE AND LIP-SYNC: Grace speaks in shot two, calm and slow. Leo reacts with surprise in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: The steady electric hum of the wheel motor. Wet slapping and a soft hiss of clay under hands. Water dripping from a sponge. A quiet studio with a light echo and distant wind chimes outside. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Calm, tactile, meditative craft video. Render with fine detail that holds up at 2K: individual threads, pores, droplets, and surface scratches should be visible on close shots. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose MiniMax H3

Native 2K output

MiniMax H3 is the only H3 tier on Fliki that renders at 2K (1440p) as well as 768p. Use 2K for hero shots, product close-ups, and anything that will be viewed on a large screen. Need faster drafts? H3 Max Turbo and H3 Fast cover that.

Sound in the same pass

Voice, sound effects, and music are generated together with the picture as one stereo track. Write dialogue in quotes and describe the ambience, and the clip arrives with audio that matches what is on screen.

Multi-shot scenes in one clip

Declare hard cuts with timecodes and H3 executes them inside a single generation, keeping the same people, wardrobe, and props across shots. It suits short ads, story beats, and explainers.

Up to 14 references

Guide a render with up to 9 reference images, 3 reference videos, and 2 reference audio clips. Lock a face, borrow a camera move, or match a voice without retraining anything.

First and last frame control

Pin an opening image, a closing image, or both, and H3 animates the path between them. It is the easiest way to keep continuity between clips in a longer Fliki project.

Long, detailed prompts

H3 accepts prompts of up to 7,000 characters and follows them closely. Write a full shot list with subject, setting, lighting, camera, dialogue, and sound, and each part is respected.

5 to 15 second clips

Choose any whole-second length from 5 to 15 seconds. Longer clips give multi-shot stories room to breathe without stitching separate generations together.

Five native aspect ratios

Render 16:9, 9:16, 1:1, 4:3, or 3:4. Each ratio is composed natively, so a vertical Reel is framed for vertical rather than cropped from widescreen.

How it works

How to generate a video with MiniMax H3

Getting a finished clip out of MiniMax H3 takes a few steps inside Fliki. Follow these six.

Fliki prompt input with a multi-shot description for the MiniMax H3 AI video generator
Step 1

Write your prompt

Open Fliki and describe the scene. Include the subject, setting, lighting, camera, any dialogue in quotes, and the sounds you want. MiniMax H3 reads up to 7,000 characters, so list shots with timecodes for multi-shot clips.

Fliki model selector with MiniMax H3 chosen for AI video generation
Step 2

Select MiniMax H3 as your model

Open the model selector and choose MiniMax H3. Fliki sends your prompt to MiniMax's model with no extra setup.

Choose an aspect ratio for MiniMax H3 video generation on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube, 9:16 for TikTok, Reels, and Shorts, 1:1 for feeds, or 4:3 and 3:4. MiniMax H3 composes each ratio natively.

Set video duration on the Fliki slider for MiniMax H3
Step 4

Set the duration

Pick any length from 5 to 15 seconds. Give multi-shot sequences 10 seconds or more so each shot has room.

Upload optional guidance media for MiniMax H3 on Fliki
Step 5

Add references (optional)

Upload reference images, video, or audio to lock a character, a camera move, or a voice, or pin a first and last frame instead. The two approaches are used separately.

Pick output resolution and generate video with MiniMax H3 on Fliki
Step 6

Select resolution and generate

Pick 768p for everyday work or 2K (1440p) for the sharpest output, then hit Generate. 2K is available on the Premium plan.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

MiniMax H3 FAQ

Frequently asked questions

Everything you need to know about generating with MiniMax H3 inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
MiniMax H3 · Free forever plan

Generate your next video with MiniMax H3.

Up to 2K resolution, native stereo audio, and multi-shot scenes from one prompt. Free to start, no credit card required.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime