video model · by MiniMax

MiniMax H3 Max AI Video Generator

Generate polished 768p videos with MiniMax H3 Max, MiniMax's performance-tuned Hailuo 3.0 tier. It takes the full reference set of 9 images, 3 videos, and 3 audio clips, follows long multi-shot prompts, and renders synced audio in clips up to 15 seconds. Compare it side by side with all our AI video models before you render.

Generated with MiniMax H3 Max

A handful of MiniMax H3 Max clips generated inside Fliki. No edits, no post.

Prompt

A 10-second 9:16 music performance clip on a rooftop in Nashville, Tennessee, at golden hour, where a singer-songwriter plays an acoustic guitar and sings a short line. THE SUBJECT: Harper, a 28-year-old white American singer with shoulder-length copper hair, a light freckled complexion, a vintage cream suede jacket with fringe, a white T-shirt, and a small turquoise ring. She plays a sunburst acoustic guitar with a woven strap. SETTING: A brick rooftop with string lights, potted plants, and a downtown skyline behind her. LIGHTING: Warm golden hour sun from camera right, flares in the lens, soft glow on her hair. CAMERA: Slow handheld orbit and close-ups, music video feel. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide: Harper sits on a stool strumming, skyline behind her. 2.5s to 5.0s: Close-up of her face, eyes closed, singing: "Carry me home tonight." 5.0s to 7.5s: Close-up of fingers on the fretboard shifting chords, strings vibrating. 7.5s to 10.0s: Medium: she opens her eyes and smiles at the lens, holding the last chord. DETAILS: Harper plays a finger-picked and strummed pattern, with her right hand strumming near the sound hole and her left hand clearly fretting chords. The guitar has a worn patch below the sound hole from years of playing. String lights hang in loose loops above her, not yet lit brightly against the sunset but glowing softly. A small potted fig tree stands at her left and a wooden crate with a vintage amp sits at her right. The Nashville skyline shows in soft focus with warm windows catching the low sun. Her jacket fringe sways gently with her movement and the breeze. Her performance is emotional and restrained: small head tilts, a slight lift of her chin on the word "home," and a breath before the line. Her fingers move in time with the strumming sound, and the strum rhythm never drifts from what is heard. The camera orbits a quarter circle around her in shot one and stays locked in shots two and three. DIALOGUE AND LIP-SYNC: Harper sings one line in shot two, gentle and breathy, in time with the strum. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Warm acoustic guitar strumming a slow G to C progression throughout, her voice clear on top, a soft breeze, faint city traffic below. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Intimate, cinematic music video. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 family cooking moment in a home kitchen in San Antonio, Texas, where a grandmother teaches her granddaughter to spread masa for tamales. THE SUBJECT: Abuela Carmen, a 76-year-old Mexican American woman with silver hair in a low bun, reading glasses on a beaded chain, and a floral apron over a blue blouse. Her granddaughter Sofia, 19, with long dark hair in a claw clip and a grey university hoodie. SETTING: A warm kitchen with yellow tiles, a big steaming pot, a stack of soaked corn husks, a bowl of masa, and a crucifix on the wall. LIGHTING: Warm afternoon window light and a soft overhead bulb, steam glowing. CAMERA: Close handheld, intimate, eye level. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Close-up: Carmen's wrinkled hands spread masa thinly on a corn husk with a spoon. 2.5s to 5.0s: Medium two-shot: Sofia tries, too thick. Carmen says: "Thinner, mija. Like this." 5.0s to 7.5s: Close-up: Sofia spreads it right, Carmen nods. 7.5s to 10.0s: Close-up on Sofia smiling. She says: "Like this, Abuela?" DETAILS: The kitchen table is covered with a plastic floral tablecloth, and the ingredients sit in fixed places: masa bowl center, corn husks soaking in a large glass bowl on the left, a pot of red chile pork filling on the right. Carmen's reading glasses sit on her nose, and her beaded chain swings slightly when she leans. Sofia's sleeves are pushed up to her elbows. A few family photos in frames sit on a shelf by the window, and a pot of basil grows on the sill. Steam rises from the big pot on the stove behind them in every shot. Carmen's performance is patient, loving, and a little bit teasing; Sofia is focused and eager to get it right. Their eye lines meet in shot two and again in shot four when Sofia looks up for approval. Carmen's hands are small, wrinkled, and confident; Sofia's are smooth and hesitant. DIALOGUE AND LIP-SYNC: Carmen speaks gently with a Spanish accent in shot two. Sofia answers in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Kitchen ambience, steam hissing from the pot, spoon scraping husk, a radio playing a faint ranchera. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Tender family story. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 skate clip at the Venice Beach skatepark in Los Angeles, California, following a skater landing a trick while his friend films. THE SUBJECT: Jalen, a 21-year-old Black American skater with short twists, a white tee, baggy khaki pants, and worn black skate shoes, riding a maple deck with yellow wheels. His friend Chris, 22, a Korean American guy in a bucket hat holding a phone. SETTING: Concrete bowls near the beach, palm trees, and the ocean in the distance. LIGHTING: Bright afternoon California sun, crisp shadows. CAMERA: Low follow-cam style, fisheye-like wide, then close-ups. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Low follow shot: Jalen pushes toward the ledge. 2.5s to 5.0s: Wide: he pops a kickflip over the gap, board spinning under him. 5.0s to 7.5s: Close-up: wheels land on concrete, he rolls away. 7.5s to 10.0s: Medium: Chris runs up with the phone. He shouts: "First try, bro!" DETAILS: Jalen's board has a scuffed black grip tape top and a colorful graphic underneath visible only during the flip, with no readable words. His white tee flutters as he rides. Chris wears a tan bucket hat, a black graphic tee without readable text, and cargo shorts, and stands on the edge of the bowl filming in shots one to three, then runs in for shot four. The concrete is sun-bleached grey with chalk marks and dark wheel scuffs. The kickflip is fast and physically believable: the board rotates exactly once along its long axis, Jalen's feet catch it on the bolts, his knees bend to absorb the landing, and his arms swing for balance. Background skaters sit on the bowl edge and a few beach walkers pass by. Palm trees sway lightly against a clear blue sky. DIALOGUE AND LIP-SYNC: Chris shouts in shot four, hyped. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Wheels rolling on concrete, the loud pop and clack of the board, distant ocean waves, onlookers cheering. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Raw skate video energy. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 real estate home tour clip in a craftsman house in Raleigh, North Carolina, where an agent welcomes viewers through the front door. THE SUBJECT: Nicole, a 36-year-old white American agent with a sleek brown bob, a camel blazer over a black top, and gold hoops, holding a set of keys. SETTING: A craftsman bungalow with a deep porch, a teal front door, hardwood floors, and an open living room with a fireplace. LIGHTING: Soft morning sun through big windows. CAMERA: Smooth gimbal walk-through. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Medium: Nicole unlocks the teal door. She says: "Come on in." 2.5s to 5.0s: Gimbal follow into the living room, sunlight on hardwood. 5.0s to 7.5s: Wide of the fireplace and built-in shelves. 7.5s to 10.0s: Medium: she turns to the lens. She says: "This one will not last." DETAILS: The porch has two white rocking chairs, a potted fern hanging from the ceiling, and brass house numbers "412" beside the door. Inside, the living room has a light oak floor, a cream sofa with linen cushions, a woven jute rug, a brick fireplace painted white with a black iron screen, and built-in shelves with books and small ceramics on both sides. Nicole moves at an easy walking pace, turning partway to invite the viewer along, and gestures with an open palm toward the fireplace. Her blazer, hair, and hoops stay identical in every shot. Her delivery is confident, friendly, and conversational, like she is showing a friend around. The gimbal move in shot two is smooth and slow, gliding forward at chest height. Sunlight streams in from tall windows on the right, casting soft window-pane shadows on the floor. DIALOGUE AND LIP-SYNC: Nicole speaks warmly in shots one and four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Keys jingling, the door creaking, footsteps on hardwood, birds outside. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Polished real estate marketing. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 9:16 quiet night shift moment in a hospital break room in Philadelphia, Pennsylvania, as two nurses share coffee at 3 a.m. THE SUBJECT: Tanya, a 44-year-old Black American nurse with braids in a bun and navy scrubs. Her colleague Josh, a 31-year-old white American nurse with a buzz cut and teal scrubs. SETTING: A small break room with a vending machine, a microwave, a round table, and a clock reading 3:04. LIGHTING: Dim fluorescent light and the glow of the vending machine. CAMERA: Static and slow push-ins. FRAMING (9:16 vertical): compose for a phone screen. Keep faces and hands in the middle two thirds of the frame, with headroom of roughly one tenth of the height. Leave the top 12 percent and bottom 18 percent free of key action, because platform buttons and captions sit there. Stack depth vertically (foreground detail low, subject center, background high) instead of spreading it sideways. Never letterbox, never pillarbox, never place a horizontal image inside the vertical frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide: Tanya slumps into a chair with a paper cup. 2.5s to 5.0s: Medium: Josh sits across and says: "Rough one tonight?" 5.0s to 7.5s: Close-up on Tanya, tired smile. She says: "We saved him, though." 7.5s to 10.0s: Medium: they tap paper cups together. DETAILS: The break room has pale green walls, a corkboard with pinned papers too small to read, a sink with a stack of mugs, and a small window showing a dark parking lot with orange sodium lights. The wall clock reads 3:04 and the hands stay consistent. Tanya's ID badge hangs from a lanyard, and she has a stethoscope around her neck. Josh has a pen behind his ear. Both look exhausted in a real way: shadowed eyes, loosened posture, slight slump. Their paper cups are plain white with brown cardboard sleeves. Tanya's smile in shot three is small and heavy with relief, and her eyes glisten slightly. Josh listens with real attention, nodding once. The final cup tap is small and quiet, a shared moment. The vending machine glows cool white on one side of their faces while the ceiling tubes give a flat greenish fill. DIALOGUE AND LIP-SYNC: Josh speaks softly in shot two, Tanya in shot three. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Vending machine hum, a distant hospital beep, a pager buzz. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Quiet human drama. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 16:9 noir scene in a 24-hour diner in Chicago, Illinois, at night, where a detective questions a nervous witness. THE SUBJECT: Detective Ray Moreno, a 50-year-old Cuban American man with a grey mustache, a rumpled trench coat, and a loosened tie. The witness Clara, a 35-year-old white American woman with red lipstick, a green raincoat, and wet hair. SETTING: A booth by a rain-streaked window, neon sign outside, coffee cups on the table. LIGHTING: Red and blue neon through the rain, a warm overhead lamp. CAMERA: Wide establishing, then over-the-shoulder close-ups. FRAMING (16:9 widescreen): compose for a landscape screen. Use the full width: place subjects on the left or right third and let the environment breathe on the other side. Keep horizons level and straight. Wide shots should show real geography so the viewer understands where everyone stands, and closer shots should keep the eye line consistent with the wide. No black bars, no split screens. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Wide: the booth by the window, rain outside, neon glow. 2.5s to 5.0s: Over the shoulder on Clara. Ray says: "You saw his face." 5.0s to 7.5s: Close-up on Clara, trembling. She says: "I saw his car." 7.5s to 10.0s: Close-up on Ray sliding a photo across the table. DETAILS: The diner booth is red vinyl with a small crack on the seat. On the table: two white ceramic mugs of black coffee, a chrome napkin dispenser, a sugar pourer, and a manila folder under Ray's hand. The photo he slides across is a glossy black and white print, face down until the last moment, then flipped face up but blurred by shallow focus. Rain streams down the window in thick rivulets, distorting the neon sign outside into red and blue smears with no readable letters. Ray is calm, heavy, and patient, with tired eyes. Clara is frightened and guarded, gripping her mug with both hands, glancing toward the window. A waitress in the far background refills a coffee pot at the counter, out of focus. Every shot keeps the same side of the booth for each character: Ray on the left, Clara on the right. DIALOGUE AND LIP-SYNC: Ray speaks low in shot two, Clara nervously in shot three. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Rain on glass, a neon buzz, clinking dishes, distant thunder. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Moody neo-noir drama. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

Prompt

A 10-second 1:1 barbershop moment in Harlem, New York, where a barber finishes a fade and hands the client a mirror. THE SUBJECT: Big Mike, a 47-year-old Black American barber with a bald head, a grey beard, and a black smock. His client Devon, 26, with a fresh fade and a white tee. SETTING: A shop with red leather chairs, mirrors, and framed photos. LIGHTING: Warm overhead lights with bright mirrors. CAMERA: Centered symmetrical shots and close-ups. FRAMING (1:1 square): compose for a square feed post. Center the subject with balanced negative space on all four sides, and keep hands, faces, and the key object inside the central 80 percent so nothing important is cut by a rounded corner crop. Favor symmetrical, graphic compositions and top-down or straight-on angles. No borders, no frames inside the frame. SHOTS (four hard cuts, total exactly 10 seconds): 0.0s to 2.5s: Close-up: clippers finish the line at the temple. 2.5s to 5.0s: Medium: Mike brushes off hair and says: "Look at you now." 5.0s to 7.5s: Close-up: Devon checks the hand mirror. 7.5s to 10.0s: Medium: Devon says: "Mike, you a legend." They dap. DETAILS: The shop has black and white checkered floor tiles, three vintage red leather chairs with chrome footrests, a long mirror along the wall, and a shelf of clippers, combs, and blue barbicide jars. Framed photos of past customers and a signed basketball jersey hang on the wall with no readable text. Big Mike's smock is black with a white towel over his shoulder, and his silver clippers have a black cord. Devon sits under a black cape that Mike whips off in shot two, sending tiny hairs flying. The fade is sharp and clean with a crisp line-up at the forehead and temples. Two other customers wait on a bench in the background, one reading a magazine, one laughing at the exchange. Mike's energy is proud and fatherly; Devon is delighted, turning his head side to side in the mirror. DIALOGUE AND LIP-SYNC: Mike speaks in shot two, Devon in shot four. Every spoken line is short enough to say at a relaxed, natural pace inside its shot, never rushed. Lips, jaw, and cheeks move in sync with each syllable, breaths land between phrases, and the speaker's mouth closes when the line ends. SOUND: Clippers buzzing, a brush swish, soft talk and laughter, old soul on a radio. All sound is recorded in the space, with natural room reverb that matches the size of the location. Dialogue sits clearly above the ambience at all times. STYLE: Warm community story. Prioritize performance: believable micro-expressions, eye focus that tracks the other person, and timing that feels acted rather than posed. PERFORMANCE: Everyone on screen behaves like a real person, not a model posing. Small natural movements between lines: a blink, a shift of weight, a glance at the other person or the object in their hands. Reactions arrive a beat after the line that causes them. Nobody looks into the lens unless the shot says they speak to camera. CONTINUITY CHECKLIST: Before each cut, match the previous shot. Same wardrobe, same hair and accessories, same props in the same hands, same side of frame for each person, same time of day and weather, same light direction. Anything that was wet, dirty, cut, poured, or moved stays that way in the following shots. AUDIO MIX: Dialogue is clean, close, and centered, recorded as if on a small lavalier mic. Ambience is steady underneath and never drops out at a cut, so the four shots feel like one continuous moment. Sound effects land exactly on the on-screen action that makes them. Keep the stereo image natural, with off-screen sounds placed on the side they come from. LENS AND COLOR: Natural, true-to-life color with gentle contrast and no heavy grading, no teal and orange push, no over-sharpening. Close shots on a 50mm to 85mm equivalent with soft background blur; wide shots on a 24mm to 35mm equivalent without distortion. Subtle film grain is fine. BACKGROUND: The world around the subject is alive but never distracting. Background people, if any, go about their own business, stay out of focus, and never look at the camera or speak intelligible words. Surfaces show honest wear: fingerprints, scuffs, crumbs, dust in the light. Nothing looks staged, brand new, or showroom clean unless the scene calls for it. RULES: - Photorealistic, filmed look. Real skin texture with pores and fine hair, real fabric weave, real reflections. No plastic skin, no waxy faces, no painterly or CGI finish. - Hands have five fingers each, with natural knuckles and nails, and grip objects with believable contact and pressure. No fused, extra, or melting fingers. - The same people wear the same clothes, hair, jewelry, and makeup in every shot. Props keep the same shape, color, labels, and position between cuts unless the action moves them. - Each shot change is a clean hard cut at the stated timecode. No morphing, no dissolves, no warping between shots. - Mouths move only when that person is speaking, and lip shapes match the words. Nobody speaks off-screen unless stated. - Motion obeys gravity and momentum. Liquids pour, cloth folds, and hair moves with weight. - No on-screen text, captions, subtitles, logos, watermarks, lower thirds, or UI overlays of any kind, unless a label is described as part of a physical prop. - No background music unless the SOUND section asks for it.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose MiniMax H3 Max

The tier Fliki trusts for music videos

H3 Max is the default renderer behind Fliki's music video workflow, where every clip has to hold a character, a location, and a beat. The same model is yours to use directly.

Full multi-modal references

Attach up to 9 reference images, 3 reference videos, and 3 reference audio clips. Keep a cast consistent, borrow a camera move, or match a voice in one render.

Synced native audio

Speech, effects, and music are generated with the picture. Put dialogue in quotes and describe the room tone, and lips and sound land together.

768p or 480p

Render final shots at 768p (1344x768 in 16:9) or save credits with 480p. For 2K, MiniMax H3 is the tier to use.

Detailed prompt following

H3 Max reads prompts of up to 7,000 characters. Subject, wardrobe, lighting, lens, blocking, and sound direction all carry through.

First and last frame guidance

Pin the opening and closing image and H3 Max animates the motion between them, which keeps clips in a longer edit continuous.

5 to 15 second clips

Choose any length from 5 to 15 seconds. Longer clips leave room for three or four shots in one generation.

Five native aspect ratios

Render 16:9, 9:16, 1:1, 4:3, or 3:4, each framed natively for its platform.

How it works

How to generate a video with MiniMax H3 Max

Getting a finished clip out of MiniMax H3 Max takes a few steps inside Fliki. Follow these six.

Fliki prompt input with a multi-shot description for the MiniMax H3 Max AI video generator
Step 1

Write your prompt

Open Fliki and describe the scene. Include the subject, setting, lighting, camera, any dialogue in quotes, and the sounds you want. MiniMax H3 Max reads up to 7,000 characters, so list shots with timecodes for multi-shot clips.

Fliki model selector with MiniMax H3 Max chosen for AI video generation
Step 2

Select MiniMax H3 Max as your model

Open the model selector and choose MiniMax H3 Max. Fliki sends your prompt to MiniMax's model with no extra setup.

Choose an aspect ratio for MiniMax H3 Max video generation on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube, 9:16 for TikTok, Reels, and Shorts, 1:1 for feeds, or 4:3 and 3:4. MiniMax H3 Max composes each ratio natively.

Set video duration on the Fliki slider for MiniMax H3 Max
Step 4

Set the duration

Pick any length from 5 to 15 seconds. Give multi-shot sequences 10 seconds or more so each shot has room.

Upload optional guidance media for MiniMax H3 Max on Fliki
Step 5

Add references (optional)

Upload reference images, video, or audio to lock a character, a camera move, or a voice, or pin a first and last frame instead. The two approaches are used separately.

Pick output resolution and generate video with MiniMax H3 Max on Fliki
Step 6

Select resolution and generate

Pick 768p for final output or 480p for a cheaper take, then hit Generate. Need 2K? Switch to MiniMax H3.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

MiniMax H3 Max FAQ

Frequently asked questions

Everything you need to know about generating with MiniMax H3 Max inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
MiniMax H3 Max · Free forever plan

Generate your next video with MiniMax H3 Max.

Full references, synced audio, and multi-shot scenes from the performance-tuned Hailuo tier. Free to start, no credit card required.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime