video model · by Pruna AI
P-Video 2 Pro AI Video Generator
Generate cinematic clips from a text prompt or a first and last frame with P-Video 2 Pro, Pruna AI's fast video model built on MiniMax H3. Render 5 to 15 second videos at 480p or 720p for a fraction of the credits of flagship models, inside Fliki. Compare it side by side with all our AI video models before you render.
Generated with P-Video 2 Pro
P-Video 2 Pro clips generated inside Fliki. No edits, no post.
Prompt
SNEAKER UNBOXING UGC. A fast, satisfying unboxing video for a new running shoe, filmed from the creator point of view on a bedroom floor. It should feel like authentic social content with tactile hands on moments and a clean final product reveal, told entirely through action and sound with no talking. THE SUBJECT: Jordan, a Latino American teenager of about nineteen, seen mostly as hands and forearms with medium tan skin, a thin silver chain bracelet on the right wrist, and the cuffs of a heather grey hoodie. In shot 3 his face appears: short curly dark hair, a faded haircut, and a delighted open mouthed grin. The product is a white running shoe with a mint green sole, a translucent heel counter, and a small black tongue tab printed with the word GLIDE in capital letters. It comes in a matte black shoebox with a lift off lid and white tissue paper inside. SETTING: A bedroom floor with light oak laminate, a corner of a grey rug, a pair of wireless headphones and a phone nearby, and a softly blurred bed with a navy comforter in the background. Casual and real, lightly messy. LIGHTING: Late afternoon daylight from a window on camera right, falling across the floor in a soft warm patch, with gentle shadows under the box. The white shoe reads bright but keeps texture in the mesh. No studio lights. CAMERA: Top down and point of view phone angles with natural slight handheld drift, quick but clean reframes. Shot 3 flips to a front facing angle at chest height. Focus snaps crisply onto the product in each shot. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Top down. Both hands lift the lid off the black shoebox and peel back the white tissue paper in one smooth motion, revealing the pair of shoes nested heel to toe. SHOT 2, 3 to 6 seconds: Close point of view. The right hand lifts one shoe out and rotates it slowly so the mint green sole, translucent heel, and the GLIDE tongue tab each catch the window light. The left thumb presses into the foam midsole and it springs back. SHOT 3, 6 to 10 seconds: Front angle. Jordan holds the shoe up next to his face, grins widely at the camera, raises his eyebrows, and gives the sole a playful double tap with two fingers, then lowers it into his lap. SOUND: Crisp cardboard friction as the lid lifts, rustling tissue paper, the soft squeak of rubber on fingertips, a hollow tap on the sole in shot 3, and a quiet room tone with faint traffic outside. A punchy lo fi hip hop beat sits underneath at low volume. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Bright, clean social media look with natural colors, true whites, gentle contrast, crisp product detail, and a subtle warm cast from the window. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
NEON NOIR DETECTIVE. A moody neo noir short: a detective in a rain soaked city alley discovers a clue under a flickering neon sign. The goal is strong cinematic atmosphere, reflections, rain physics, and a consistent character across three angles, with no dialogue. THE SUBJECT: Detective Elena Park, a Korean American woman in her forties with a sharp chin length black bob, a serious focused face, and faint laugh lines. She wears a long tan trench coat with the collar turned up, black leather gloves, and dark trousers. She carries a small silver flashlight. The clue is a single red playing card, the queen of hearts, lying face up in a puddle. SETTING: A narrow downtown alley at night after heavy rain. Wet brick walls, fire escapes, steam drifting from a grate, overflowing trash cans, and a pink and cyan neon sign above a back door that reads NIGHT OWL in cursive script, one letter flickering. Puddles reflect everything. LIGHTING: Hard pink and cyan neon from above mixed with a cool blue moonlight fill, deep black shadows. The flashlight throws a tight white beam through the rain. Raindrops catch the light like thin silver streaks. CAMERA: Cinematic 35mm look with anamorphic flares from the neon. Shot 1 is a slow low tracking shot behind her. Shot 2 is a tight overhead insert. Shot 3 is a slow push in on her face. Deliberate, weighty moves. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Low angle from behind as Elena walks down the alley toward the neon sign, trench coat swaying, boots splashing through a puddle that ripples the neon reflection. Rain falls steadily. SHOT 2, 3 to 6 seconds: Overhead insert. Her gloved hand enters frame and the flashlight beam lands on the red queen of hearts in the puddle, the card rippling as raindrops hit the water around it. SHOT 3, 6 to 10 seconds: Slow push in on Elena crouching, rain dripping from her hair, the flickering neon washing her face pink and cyan. Her eyes narrow and she looks up slowly toward the fire escape above, lips pressed together. SOUND: Steady rain on brick and metal, heavy drips from a fire escape, the electric buzz and tick of the flickering neon, boots splashing through puddles, a distant police siren, and a low moody saxophone and upright bass under the scene. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Neo noir grade with saturated pink and cyan highlights, deep inky blacks, glossy wet surfaces, fine grain, and soft bloom on the neon. It should look like a frame from a stylish crime thriller. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
PERFUME BOTTLE REVEAL. A luxury fragrance product film where a glass perfume bottle is revealed through flowing silk and water. The priority is elegant slow motion, liquid and fabric physics, and immaculate glass rendering. THE SUBJECT: A heavy faceted glass perfume bottle, roughly square with beveled edges, filled with pale amber liquid, topped with a polished gold cap shaped like a smooth pebble. The front of the glass is etched with the single word LUMEN in thin elegant capital letters. A length of champagne colored silk fabric and a few floating white jasmine flowers complete the scene. The bottle is identical in every shot. SETTING: An abstract studio set: a shallow pool of still black water on a black stone plinth, with a dark gradient background that falls to pure black. Nothing else in frame. LIGHTING: A single hard backlight from behind and above that makes the amber liquid glow and the facets sparkle, with a soft strip light on the right carving a clean highlight down the bottle edge. Caustic light patterns ripple on the plinth from the water. CAMERA: High end product cinematography on a 100mm macro lens. Slow motion feel with smooth motion control moves. Shot 1 is a macro drift, shot 2 a slow orbit, shot 3 a locked hero frame. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Macro. Champagne silk slides slowly across the frame, its folds rippling, and pulls away to reveal the gold cap and the top facets of the bottle, light glinting as it passes. SHOT 2, 3 to 7 seconds: Slow orbit around the bottle standing in the shallow water. A single jasmine flower drops into the pool beside it and sends out a perfect ring of ripples. Caustics dance across the bottle and plinth. SHOT 3, 7 to 10 seconds: Locked hero frame. The bottle stands centered, the etched LUMEN clearly readable, amber liquid glowing, jasmine flowers drifting slowly on the dark water around its base as the last ripple fades. SOUND: A whisper of silk sliding, a single delicate water drop, soft ripples, a faint glassy shimmer, and a slow warm ambient synth pad with a gentle piano note on the hero frame. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Luxury commercial grade: deep blacks, rich amber and gold, champagne highlights, silky smooth gradients, and flawless glass with no warping or artifacts. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
BOXING GYM TRAINING. An intense sports clip of a boxer training on a heavy bag in an old gym, built for vertical social feeds. It should show power, sweat, speed, and body mechanics in three punchy shots. THE SUBJECT: Tasha, a Black American woman in her late twenties with an athletic muscular build, her hair in tight cornrow braids pulled back, and a determined face with a thin sheen of sweat. She wears a black sports bra, red high waisted training shorts, white hand wraps, and red leather boxing gloves. Her wraps, gloves, and outfit are identical in all shots. SETTING: A gritty old boxing gym in Philadelphia: worn red canvas heavy bags hanging on chains, a boxing ring with frayed ropes in the background, peeling posters on brick walls, dusty wood floors, and a big round wall clock. LIGHTING: Hard overhead industrial lights in cages creating pools of light and dark, with a dusty shaft of daylight from a high window catching floating particles. Sweat catches small highlights on her shoulders and face. CAMERA: Sports documentary style with a mix of handheld and speed ramps. Shot 1 is a medium tracking shot, shot 2 a low slow motion close up, shot 3 a handheld close up. Fast, energetic, but readable. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Medium shot circling Tasha as she throws a fast jab, jab, cross combination into the heavy bag, the bag swinging and the chain rattling. Her feet pivot correctly on each punch. SHOT 2, 3 to 7 seconds: Low angle slow motion close up. Her right hook lands deep in the bag, sweat droplets spray off her glove and shoulder into the light shaft, dust shakes loose from the bag surface, the canvas visibly deforms and rebounds. SHOT 3, 7 to 10 seconds: Handheld close up. Tasha steps back, breathing hard, tucks her gloves to her chin, and stares straight into the lens with total focus before snapping one more jab toward the camera. SOUND: Heavy thuds of gloves on leather, the rattle and creak of chains, the squeak of shoes on the wood floor, sharp exhales on each punch, a round timer bell far off, and a driving hip hop drum beat with deep bass. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: High contrast sports grade: warm skin tones, deep shadows, gritty textures, and punchy saturation on the red gloves and bags. It should look like a premium athletic brand spot. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
KYOTO MORNING TRAVEL. A calm travel film of early morning in Kyoto following one traveler through a quiet lane to a temple gate. It should feel peaceful and cinematic with real atmosphere and graceful camera movement. THE SUBJECT: Olivia, a white American woman in her early thirties with shoulder length light brown hair, wearing a long camel wool coat, a cream turtleneck, dark jeans, and white sneakers, carrying a small brown leather crossbody bag and holding a paper cup of tea. Her outfit and bag stay identical in every shot. SETTING: The Higashiyama district of Kyoto at dawn: a narrow stone paved lane with traditional wooden machiya houses, paper lanterns hanging from eaves, and a five story wooden pagoda visible at the top of the slope. Later, a vivid vermilion torii gate with a stone path beyond. The streets are nearly empty. LIGHTING: Soft blue dawn light turning pale gold as the sun rises, low mist in the air, lanterns still glowing warmly. Gentle long shadows on the stones. Delicate, luminous, never harsh. CAMERA: Smooth gimbal travel film look on a 35mm lens with gentle forward and lateral moves. Shot 1 follows from behind, shot 2 is a slow side dolly, shot 3 is a slow tilt up. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: From behind, Olivia walks slowly up the stone lane toward the pagoda, sipping her tea, the lanterns glowing on both sides and mist drifting between the houses. SHOT 2, 3 to 6 seconds: Side dolly at waist height as she passes a wooden shopfront where an older shopkeeper slides open a paper screen door. Olivia turns her head and gives a small polite bow. SHOT 3, 6 to 10 seconds: Olivia stops at the foot of the vermilion torii gate, looking up. The camera tilts up slowly past the gate to the brightening sky as the first sun rays hit the top beam. SOUND: Soft footsteps on stone, distant temple bell ringing once, birds waking up, the wooden rattle of the sliding screen door, gentle wind in the trees, and a quiet koto and ambient pad melody. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Soft, airy travel grade with pastel blues and warm golds, lifted shadows, fine detail, and a calm, poetic mood. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
RAMEN KITCHEN CLOSE UPS. A widescreen food film of a chef assembling a bowl of tonkotsu ramen in a small restaurant kitchen. The focus is steam, noodle physics, broth pour, and precise placement of toppings. THE SUBJECT: Kenji, a Japanese American chef in his fifties with short salt and pepper hair under a navy blue head wrap, wearing a black chef jacket with the sleeves rolled. His hands are strong and precise. The bowl is a wide black ceramic bowl with a thin red rim. Toppings: two slices of chashu pork, a halved soft boiled egg with a jammy orange yolk, green onions, a sheet of nori, and a few slices of pink and white naruto fish cake. SETTING: A narrow ramen shop kitchen with stainless steel counters, big stock pots steaming on burners, a noodle basket station, wooden shelves with bowls, and a short cloth curtain at the pass. Warm and busy but tidy. A row of red paper lanterns hangs over the counter seating just beyond the pass, softly out of focus, and a handwritten wooden menu board hangs on the back wall without readable text. LIGHTING: Warm overhead pendant lights and a strong backlight through the steam so it glows. The broth surface shows glossy highlights and small fat droplets. Rich, appetizing contrast. CAMERA: Food cinematography on 50mm and 100mm lenses with slow dolly moves and one overhead shot. Crisp focus on the food with a shallow background. FRAMING (16:9 landscape): Compose for a widescreen display. Use the full width for environment and depth, place subjects on rule of thirds lines, and let the background carry story information. Wide and medium shots should breathe; close ups should still show some of the environment at the edges. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Medium shot. Kenji lifts a basket of noodles from boiling water and shakes it hard three times, water spraying and steam billowing in the backlight. SHOT 2, 3 to 6 seconds: Close up. Creamy white broth pours from a ladle into the black bowl over the noodles in a smooth ribbon, the noodles swirling and settling. SHOT 3, 6 to 10 seconds: Overhead. His hands place the chashu, the halved egg, the naruto, and the nori one by one with chopsticks, then sprinkle green onions. The finished bowl sits centered, steam rising. SOUND: Rolling boil, the rapid shake and slap of the noodle basket, splashing water, a thick liquid pour, chopsticks clicking on ceramic, burners roaring softly, and quiet kitchen chatter in the distance. No voices in front. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Warm, rich food grade: creamy whites, deep blacks, glossy highlights, vivid green onion and orange yolk. Detailed and mouthwatering. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
ANIME ROOFTOP SUNSET. A hand drawn anime style scene of two high school friends sharing a quiet sunset on a school rooftop. The model should keep a consistent 2D cel shaded style, expressive faces, and painterly backgrounds, with gentle motion and no dialogue. THE SUBJECT: Aiko, a girl of about sixteen with long straight black hair and a blue hair clip, large expressive brown eyes, wearing a navy sailor style school uniform with a red neckerchief. Ren, a boy of about sixteen with messy chestnut hair and a small bandage on his cheek, wearing a white shirt with a loosened grey tie and a navy blazer. Their designs, hair, and uniforms are identical in every shot. SETTING: A school rooftop surrounded by a tall green chain link fence, a water tank, and a view of a seaside town and the ocean beyond, with a train line crossing a bridge in the distance. Painted watercolor clouds fill a huge sky. A pair of school bags rests against the water tank, one navy with a small yellow star charm, and a few cherry blossom petals blow across the concrete floor. LIGHTING: Golden orange sunset with long purple shadows, soft glowing rim light on the characters' hair, and lens light leaks. Sky gradient from peach to deep lavender. CAMERA: Anime cinematography: shot 1 is a wide establishing shot with slight parallax, shot 2 a close up, shot 3 a slow pull back. Movements are gentle and composed. FRAMING (1:1 square): Compose for a square feed post. Center the subject with balanced space on all sides, keep the key action inside the middle of the frame, and avoid placing anything important near the corners, which can be cropped in previews. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Wide shot. Aiko and Ren lean on the fence side by side, their hair and her neckerchief fluttering in the breeze, a train crossing the distant bridge. SHOT 2, 3 to 6 seconds: Close up on Aiko. She turns her head toward Ren, a soft blush appears, and she smiles slightly while her hair blows across her face. SHOT 3, 6 to 10 seconds: Slow pull back as Ren grins and holds out one earbud to her. She takes it, and the two of them look out at the ocean as the sun touches the horizon. SOUND: A gentle breeze, the chain link fence rattling softly, distant train clatter, seagulls, cicadas fading into evening, and a warm, nostalgic piano and strings melody. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Classic hand drawn 2D anime look with clean line art, cel shading, soft painted backgrounds, and warm nostalgic colors. Not 3D, not photoreal. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
SNEAKER UNBOXING UGC. A fast, satisfying unboxing video for a new running shoe, filmed from the creator point of view on a bedroom floor. It should feel like authentic social content with tactile hands on moments and a clean final product reveal, told entirely through action and sound with no talking. THE SUBJECT: Jordan, a Latino American teenager of about nineteen, seen mostly as hands and forearms with medium tan skin, a thin silver chain bracelet on the right wrist, and the cuffs of a heather grey hoodie. In shot 3 his face appears: short curly dark hair, a faded haircut, and a delighted open mouthed grin. The product is a white running shoe with a mint green sole, a translucent heel counter, and a small black tongue tab printed with the word GLIDE in capital letters. It comes in a matte black shoebox with a lift off lid and white tissue paper inside. SETTING: A bedroom floor with light oak laminate, a corner of a grey rug, a pair of wireless headphones and a phone nearby, and a softly blurred bed with a navy comforter in the background. Casual and real, lightly messy. LIGHTING: Late afternoon daylight from a window on camera right, falling across the floor in a soft warm patch, with gentle shadows under the box. The white shoe reads bright but keeps texture in the mesh. No studio lights. CAMERA: Top down and point of view phone angles with natural slight handheld drift, quick but clean reframes. Shot 3 flips to a front facing angle at chest height. Focus snaps crisply onto the product in each shot. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Top down. Both hands lift the lid off the black shoebox and peel back the white tissue paper in one smooth motion, revealing the pair of shoes nested heel to toe. SHOT 2, 3 to 6 seconds: Close point of view. The right hand lifts one shoe out and rotates it slowly so the mint green sole, translucent heel, and the GLIDE tongue tab each catch the window light. The left thumb presses into the foam midsole and it springs back. SHOT 3, 6 to 10 seconds: Front angle. Jordan holds the shoe up next to his face, grins widely at the camera, raises his eyebrows, and gives the sole a playful double tap with two fingers, then lowers it into his lap. SOUND: Crisp cardboard friction as the lid lifts, rustling tissue paper, the soft squeak of rubber on fingertips, a hollow tap on the sole in shot 3, and a quiet room tone with faint traffic outside. A punchy lo fi hip hop beat sits underneath at low volume. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Bright, clean social media look with natural colors, true whites, gentle contrast, crisp product detail, and a subtle warm cast from the window. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
NEON NOIR DETECTIVE. A moody neo noir short: a detective in a rain soaked city alley discovers a clue under a flickering neon sign. The goal is strong cinematic atmosphere, reflections, rain physics, and a consistent character across three angles, with no dialogue. THE SUBJECT: Detective Elena Park, a Korean American woman in her forties with a sharp chin length black bob, a serious focused face, and faint laugh lines. She wears a long tan trench coat with the collar turned up, black leather gloves, and dark trousers. She carries a small silver flashlight. The clue is a single red playing card, the queen of hearts, lying face up in a puddle. SETTING: A narrow downtown alley at night after heavy rain. Wet brick walls, fire escapes, steam drifting from a grate, overflowing trash cans, and a pink and cyan neon sign above a back door that reads NIGHT OWL in cursive script, one letter flickering. Puddles reflect everything. LIGHTING: Hard pink and cyan neon from above mixed with a cool blue moonlight fill, deep black shadows. The flashlight throws a tight white beam through the rain. Raindrops catch the light like thin silver streaks. CAMERA: Cinematic 35mm look with anamorphic flares from the neon. Shot 1 is a slow low tracking shot behind her. Shot 2 is a tight overhead insert. Shot 3 is a slow push in on her face. Deliberate, weighty moves. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Low angle from behind as Elena walks down the alley toward the neon sign, trench coat swaying, boots splashing through a puddle that ripples the neon reflection. Rain falls steadily. SHOT 2, 3 to 6 seconds: Overhead insert. Her gloved hand enters frame and the flashlight beam lands on the red queen of hearts in the puddle, the card rippling as raindrops hit the water around it. SHOT 3, 6 to 10 seconds: Slow push in on Elena crouching, rain dripping from her hair, the flickering neon washing her face pink and cyan. Her eyes narrow and she looks up slowly toward the fire escape above, lips pressed together. SOUND: Steady rain on brick and metal, heavy drips from a fire escape, the electric buzz and tick of the flickering neon, boots splashing through puddles, a distant police siren, and a low moody saxophone and upright bass under the scene. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Neo noir grade with saturated pink and cyan highlights, deep inky blacks, glossy wet surfaces, fine grain, and soft bloom on the neon. It should look like a frame from a stylish crime thriller. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
RAMEN KITCHEN CLOSE UPS. A widescreen food film of a chef assembling a bowl of tonkotsu ramen in a small restaurant kitchen. The focus is steam, noodle physics, broth pour, and precise placement of toppings. THE SUBJECT: Kenji, a Japanese American chef in his fifties with short salt and pepper hair under a navy blue head wrap, wearing a black chef jacket with the sleeves rolled. His hands are strong and precise. The bowl is a wide black ceramic bowl with a thin red rim. Toppings: two slices of chashu pork, a halved soft boiled egg with a jammy orange yolk, green onions, a sheet of nori, and a few slices of pink and white naruto fish cake. SETTING: A narrow ramen shop kitchen with stainless steel counters, big stock pots steaming on burners, a noodle basket station, wooden shelves with bowls, and a short cloth curtain at the pass. Warm and busy but tidy. A row of red paper lanterns hangs over the counter seating just beyond the pass, softly out of focus, and a handwritten wooden menu board hangs on the back wall without readable text. LIGHTING: Warm overhead pendant lights and a strong backlight through the steam so it glows. The broth surface shows glossy highlights and small fat droplets. Rich, appetizing contrast. CAMERA: Food cinematography on 50mm and 100mm lenses with slow dolly moves and one overhead shot. Crisp focus on the food with a shallow background. FRAMING (16:9 landscape): Compose for a widescreen display. Use the full width for environment and depth, place subjects on rule of thirds lines, and let the background carry story information. Wide and medium shots should breathe; close ups should still show some of the environment at the edges. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Medium shot. Kenji lifts a basket of noodles from boiling water and shakes it hard three times, water spraying and steam billowing in the backlight. SHOT 2, 3 to 6 seconds: Close up. Creamy white broth pours from a ladle into the black bowl over the noodles in a smooth ribbon, the noodles swirling and settling. SHOT 3, 6 to 10 seconds: Overhead. His hands place the chashu, the halved egg, the naruto, and the nori one by one with chopsticks, then sprinkle green onions. The finished bowl sits centered, steam rising. SOUND: Rolling boil, the rapid shake and slap of the noodle basket, splashing water, a thick liquid pour, chopsticks clicking on ceramic, burners roaring softly, and quiet kitchen chatter in the distance. No voices in front. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Warm, rich food grade: creamy whites, deep blacks, glossy highlights, vivid green onion and orange yolk. Detailed and mouthwatering. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
PERFUME BOTTLE REVEAL. A luxury fragrance product film where a glass perfume bottle is revealed through flowing silk and water. The priority is elegant slow motion, liquid and fabric physics, and immaculate glass rendering. THE SUBJECT: A heavy faceted glass perfume bottle, roughly square with beveled edges, filled with pale amber liquid, topped with a polished gold cap shaped like a smooth pebble. The front of the glass is etched with the single word LUMEN in thin elegant capital letters. A length of champagne colored silk fabric and a few floating white jasmine flowers complete the scene. The bottle is identical in every shot. SETTING: An abstract studio set: a shallow pool of still black water on a black stone plinth, with a dark gradient background that falls to pure black. Nothing else in frame. LIGHTING: A single hard backlight from behind and above that makes the amber liquid glow and the facets sparkle, with a soft strip light on the right carving a clean highlight down the bottle edge. Caustic light patterns ripple on the plinth from the water. CAMERA: High end product cinematography on a 100mm macro lens. Slow motion feel with smooth motion control moves. Shot 1 is a macro drift, shot 2 a slow orbit, shot 3 a locked hero frame. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Macro. Champagne silk slides slowly across the frame, its folds rippling, and pulls away to reveal the gold cap and the top facets of the bottle, light glinting as it passes. SHOT 2, 3 to 7 seconds: Slow orbit around the bottle standing in the shallow water. A single jasmine flower drops into the pool beside it and sends out a perfect ring of ripples. Caustics dance across the bottle and plinth. SHOT 3, 7 to 10 seconds: Locked hero frame. The bottle stands centered, the etched LUMEN clearly readable, amber liquid glowing, jasmine flowers drifting slowly on the dark water around its base as the last ripple fades. SOUND: A whisper of silk sliding, a single delicate water drop, soft ripples, a faint glassy shimmer, and a slow warm ambient synth pad with a gentle piano note on the hero frame. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Luxury commercial grade: deep blacks, rich amber and gold, champagne highlights, silky smooth gradients, and flawless glass with no warping or artifacts. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
ANIME ROOFTOP SUNSET. A hand drawn anime style scene of two high school friends sharing a quiet sunset on a school rooftop. The model should keep a consistent 2D cel shaded style, expressive faces, and painterly backgrounds, with gentle motion and no dialogue. THE SUBJECT: Aiko, a girl of about sixteen with long straight black hair and a blue hair clip, large expressive brown eyes, wearing a navy sailor style school uniform with a red neckerchief. Ren, a boy of about sixteen with messy chestnut hair and a small bandage on his cheek, wearing a white shirt with a loosened grey tie and a navy blazer. Their designs, hair, and uniforms are identical in every shot. SETTING: A school rooftop surrounded by a tall green chain link fence, a water tank, and a view of a seaside town and the ocean beyond, with a train line crossing a bridge in the distance. Painted watercolor clouds fill a huge sky. A pair of school bags rests against the water tank, one navy with a small yellow star charm, and a few cherry blossom petals blow across the concrete floor. LIGHTING: Golden orange sunset with long purple shadows, soft glowing rim light on the characters' hair, and lens light leaks. Sky gradient from peach to deep lavender. CAMERA: Anime cinematography: shot 1 is a wide establishing shot with slight parallax, shot 2 a close up, shot 3 a slow pull back. Movements are gentle and composed. FRAMING (1:1 square): Compose for a square feed post. Center the subject with balanced space on all sides, keep the key action inside the middle of the frame, and avoid placing anything important near the corners, which can be cropped in previews. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Wide shot. Aiko and Ren lean on the fence side by side, their hair and her neckerchief fluttering in the breeze, a train crossing the distant bridge. SHOT 2, 3 to 6 seconds: Close up on Aiko. She turns her head toward Ren, a soft blush appears, and she smiles slightly while her hair blows across her face. SHOT 3, 6 to 10 seconds: Slow pull back as Ren grins and holds out one earbud to her. She takes it, and the two of them look out at the ocean as the sun touches the horizon. SOUND: A gentle breeze, the chain link fence rattling softly, distant train clatter, seagulls, cicadas fading into evening, and a warm, nostalgic piano and strings melody. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Classic hand drawn 2D anime look with clean line art, cel shading, soft painted backgrounds, and warm nostalgic colors. Not 3D, not photoreal. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
BOXING GYM TRAINING. An intense sports clip of a boxer training on a heavy bag in an old gym, built for vertical social feeds. It should show power, sweat, speed, and body mechanics in three punchy shots. THE SUBJECT: Tasha, a Black American woman in her late twenties with an athletic muscular build, her hair in tight cornrow braids pulled back, and a determined face with a thin sheen of sweat. She wears a black sports bra, red high waisted training shorts, white hand wraps, and red leather boxing gloves. Her wraps, gloves, and outfit are identical in all shots. SETTING: A gritty old boxing gym in Philadelphia: worn red canvas heavy bags hanging on chains, a boxing ring with frayed ropes in the background, peeling posters on brick walls, dusty wood floors, and a big round wall clock. LIGHTING: Hard overhead industrial lights in cages creating pools of light and dark, with a dusty shaft of daylight from a high window catching floating particles. Sweat catches small highlights on her shoulders and face. CAMERA: Sports documentary style with a mix of handheld and speed ramps. Shot 1 is a medium tracking shot, shot 2 a low slow motion close up, shot 3 a handheld close up. Fast, energetic, but readable. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Medium shot circling Tasha as she throws a fast jab, jab, cross combination into the heavy bag, the bag swinging and the chain rattling. Her feet pivot correctly on each punch. SHOT 2, 3 to 7 seconds: Low angle slow motion close up. Her right hook lands deep in the bag, sweat droplets spray off her glove and shoulder into the light shaft, dust shakes loose from the bag surface, the canvas visibly deforms and rebounds. SHOT 3, 7 to 10 seconds: Handheld close up. Tasha steps back, breathing hard, tucks her gloves to her chin, and stares straight into the lens with total focus before snapping one more jab toward the camera. SOUND: Heavy thuds of gloves on leather, the rattle and creak of chains, the squeak of shoes on the wood floor, sharp exhales on each punch, a round timer bell far off, and a driving hip hop drum beat with deep bass. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: High contrast sports grade: warm skin tones, deep shadows, gritty textures, and punchy saturation on the red gloves and bags. It should look like a premium athletic brand spot. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Prompt
KYOTO MORNING TRAVEL. A calm travel film of early morning in Kyoto following one traveler through a quiet lane to a temple gate. It should feel peaceful and cinematic with real atmosphere and graceful camera movement. THE SUBJECT: Olivia, a white American woman in her early thirties with shoulder length light brown hair, wearing a long camel wool coat, a cream turtleneck, dark jeans, and white sneakers, carrying a small brown leather crossbody bag and holding a paper cup of tea. Her outfit and bag stay identical in every shot. SETTING: The Higashiyama district of Kyoto at dawn: a narrow stone paved lane with traditional wooden machiya houses, paper lanterns hanging from eaves, and a five story wooden pagoda visible at the top of the slope. Later, a vivid vermilion torii gate with a stone path beyond. The streets are nearly empty. LIGHTING: Soft blue dawn light turning pale gold as the sun rises, low mist in the air, lanterns still glowing warmly. Gentle long shadows on the stones. Delicate, luminous, never harsh. CAMERA: Smooth gimbal travel film look on a 35mm lens with gentle forward and lateral moves. Shot 1 follows from behind, shot 2 is a slow side dolly, shot 3 is a slow tilt up. FRAMING (9:16 vertical): Compose for a phone screen held upright. Keep faces and the key action in the center two thirds of the frame, with the eyes roughly one third down from the top. Leave clear headroom and keep important detail away from the bottom fifth, where app interface overlays sit. Favor medium close ups and vertical depth over wide horizontal spreads. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: From behind, Olivia walks slowly up the stone lane toward the pagoda, sipping her tea, the lanterns glowing on both sides and mist drifting between the houses. SHOT 2, 3 to 6 seconds: Side dolly at waist height as she passes a wooden shopfront where an older shopkeeper slides open a paper screen door. Olivia turns her head and gives a small polite bow. SHOT 3, 6 to 10 seconds: Olivia stops at the foot of the vermilion torii gate, looking up. The camera tilts up slowly past the gate to the brightening sky as the first sun rays hit the top beam. SOUND: Soft footsteps on stone, distant temple bell ringing once, birds waking up, the wooden rattle of the sliding screen door, gentle wind in the trees, and a quiet koto and ambient pad melody. No voices. MOTION AND PACING: This clip tells its story through action, not speech. Nobody speaks, and any mouths on screen stay closed or move only with breathing and expression. Give every shot one clear primary action that begins right after the cut and completes before the next cut, so each beat reads cleanly on a phone. Secondary motion supports the main action: hair and fabric respond to movement and air, background elements move gently and never pull focus, and particles such as steam, dust, rain, or spray drift with believable speed and direction. Camera movement is motivated and steady, with no sudden zooms or warping. The energy builds from shot 1 to shot 3, and the final second rests on a clean, well composed frame that could work as a thumbnail. STYLE AND GRADE: Soft, airy travel grade with pastel blues and warm golds, lifted shadows, fine detail, and a calm, poetic mood. PRIORITIES: Read the whole prompt as one continuous scene with exact timings. First priority is continuity: the same faces, the same hair, the same wardrobe, the same props, and the same time of day in every shot, with lighting direction that stays consistent across cuts. Second priority is believable motion with correct weight, contact, and timing, so every action starts and finishes inside its shot. Third priority is clean synchronized audio where every sound lines up with the action that causes it. When in doubt, keep it simple, grounded, and specific rather than adding extra elements that were not requested. RULES: - Photorealistic, natural footage unless the style section says otherwise. Real skin texture with pores and fine variation, no plastic smoothing, no uncanny eyes. - Hands are anatomically correct: five fingers per hand, natural joints, correct grip on every object, no merging of fingers with props. - Characters, wardrobe, hair, props, and colors stay identical in every shot. Nothing appears, disappears, or changes color between cuts unless the shot list says so. - Motion follows real physics: weight, momentum, cloth drape, liquid behavior, and contact shadows all behave the way they do in real camera footage. - No captions, no subtitles, no on-screen titles, no logos, no watermarks, no user interface graphics. Any text that appears on a physical object is spelled exactly as written in this prompt and stays legible and stable. - Cuts happen exactly at the listed timecodes. Each shot is a clean cut, not a morph or a dissolve, unless a transition is specified. - Audio is clean and mixed like a finished piece: dialogue or the main sound sits on top, ambience underneath, no clipping, no distorted music, no random extra voices.
Why creators choose P-Video 2 Pro
Text-to-video, no image needed
Unlike the original P-Video, which animates an uploaded image, P-Video 2 Pro generates whole scenes from a written prompt. Add an image only when you want it.
Built on MiniMax H3
Pruna AI based P-Video 2 Pro on MiniMax H3, a multimodal video model, and optimized it for fast inference. You get cinematic motion without flagship wait times.
Fast turnaround
Pruna reports roughly 2 seconds of compute for a 5 second 480p clip in speed mode. That makes it practical to try several versions of an idea before committing.
Low credit cost
P-Video 2 Pro is one of the lowest-credit video models on Fliki, which suits drafts, storyboards, and high-volume social content.
Clips from 5 to 15 seconds
Set any duration from 5 to 15 seconds on Fliki. There is room for a short scene with a few shots, not just a single motion.
First and last frame control
Upload a start frame, an end frame, or both. The model moves between them, which is useful for product reveals and before and after transitions.
480p or 720p output
Draft at 480p, then render the version you like at 720p (Pruna calls this tier 768p). Both are available on Fliki at 24 fps.
Every social format
Compose natively for 16:9, 9:16, and 1:1 so the same idea works on YouTube, TikTok, Reels, Shorts, and feed posts.
How it works
How to generate a video with P-Video 2 Pro
Getting a P-Video 2 Pro clip out of Fliki takes a few quick steps.

Write your prompt
Open Fliki and describe the video like a shot list: subject, setting, action, camera, lighting, and mood. P-Video 2 Pro works from text alone, so no source image is needed.

Select P-Video 2 Pro as your model
Open the model selector and choose P-Video 2 Pro. Fliki sends your prompt to the model with no extra setup.

Pick your aspect ratio
Choose 16:9 for YouTube and web, 9:16 for TikTok, Reels, and Shorts, or 1:1 for square feeds.

Set the duration
Choose any length from 5 to 15 seconds. Short clips return fastest; longer ones fit a full scene.

Add a first or last frame (optional)
Upload a starting image, an ending image, or both to control where the clip begins and lands.

Select resolution and generate
Pick 480p for quick drafts or 720p for delivery, then hit Generate. Download the clip or place it into a longer Fliki project.
AI MODEL GALLERY
Built on the best AI models - ready inside Fliki
Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.
















P-Video 2 Pro FAQ
Frequently asked questions
Everything you need to know about generating with P-Video 2 Pro inside Fliki.
P-Video 2 Pro is a video generation model from Pruna AI based on MiniMax H3. It creates cinematic clips from text, a first frame, or a first and last frame, optimized for speed and low cost.
P-Video on Fliki is used to animate a portrait or product still into a talking avatar. P-Video 2 Pro is a general text-to-video model that builds full scenes from a prompt, with optional first and last frame control.
Pruna AI launched P-Video 2 Pro in September 2026.
P-Video 2 Pro is available on all paid Fliki plans, including Basic. It is not included on the Free plan.
On Fliki you can choose any duration from 5 to 15 seconds.
Fliki offers 480p and 720p in 16:9, 9:16, and 1:1, at 24 fps. Pruna labels its top tier 768p.
Yes. Upload a first frame, a last frame, or both, and the model animates between them following your prompt.
Pruna recommends covering subject, action, scene, camera, lighting, style, and audio. Be concrete about each, and break longer clips into timed shots.
Still curious?
Try Fliki free in your browser, no credit card required.
Start free→More from Fliki
AI video models
Discover more
Tools
Discover features
Generate your next video with P-Video 2 Pro.
Fast, low-credit cinematic clips from Pruna AI. Start free on Fliki, then upgrade to render.
Generate your first video freeFree forever plan · No credit card required · Cancel anytime

