video model · by Black Forest Labs

FLUX 3 Video AI Generator

Generate videos up to 20 seconds long with native dialogue, sound effects, and ambient audio using FLUX 3 Video, the multimodal model from Black Forest Labs. Write a multi-shot prompt, add an optional first or last frame, and render in 720p or 1080p inside Fliki. Compare it side by side with all our AI video models before you render.

Generated with FLUX 3 Video

FLUX 3 Video clips generated inside Fliki, with their native audio. No edits, no post.

Prompt

SKINCARE UGC AD. A handheld, selfie style user generated ad for a hydrating face serum, filmed in a real bathroom by the creator herself. THE SUBJECT: Maya, a Black American woman in her late twenties with a short natural coily afro, warm deep brown skin with a few small freckles across the nose, and a relaxed friendly face. SETTING: A small, tidy apartment bathroom in the morning. LIGHTING: Soft morning daylight from a frosted window on camera left, wrapping gently across her face. CAMERA: Phone camera energy: the lens is at arm length, slightly above eye level, with gentle natural hand shake and small reframes as she moves. FRAMING: vertical 9:16, faces and action in the center two thirds, headroom kept, nothing key in the bottom fifth. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Selfie medium close up. Maya leans into frame, holds the DEWDROP SERUM bottle up beside her cheek with the label facing the lens, raises her eyebrows, and speaks her opening line directly into the camera with a half smile. SHOT 2, 3 to 6 seconds: Tight close up of her right hand. She squeezes the dropper bulb, lifts it, and lets two clear drops fall onto her fingertips. The liquid is slightly viscous and catches the window light. She presses the fingertips gently onto her cheekbone. Her wardrobe and earrings stay the same. SHOT 3, 6 to 10 seconds: Back to the selfie angle. Maya turns her face slightly toward the window so the glow on her cheek catches the light, then looks back at the lens and says her closing line, ending with a small satisfied nod and a relaxed smile. She holds the bottle low in the frame, label still readable. SOUND: Quiet bathroom room tone with a faint echo off the tile, a soft glass clink as the dropper lifts, a tiny squish of the bulb, the light tap of fingertips on skin, and distant muffled street sound through the window. No music bed; the ad lives on her voice. DIALOGUE AND LIP SYNC: Maya speaks in a warm, casual, slightly conspiratorial tone, like telling a friend a secret, at a natural pace. In SHOT 1 her mouth is fully visible facing camera and she says: "Okay, this serum fixed my dry winter skin." In SHOT 2 she is silent. In SHOT 3, facing the lens with her mouth clearly visible, she says: "Two drops. That is it." Her lips, teeth, and jaw sync exactly to each word, with small natural breaths between phrases. STYLE: Clean, honest, natural color. Neutral white balance leaning very slightly warm, true to life skin tones, gentle contrast, no heavy filter. It should look like modern phone footage with good light. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

Prompt

ROOFTOP CINEMATIC SHORT. A cinematic two person moment on a city rooftop at dusk: a reunion between two old friends, told in three shots with a single emotional exchange. The goal is film quality performance, precise lip sync, and continuity across angles. THE SUBJECT: Daniel, a Korean American man in his early thirties with short black hair pushed back and light stubble, wearing a charcoal wool overcoat over a black crew neck sweater. SETTING: A flat rooftop above a Chicago apartment building, gravel surface, a low brick parapet, a rusted water tower silhouette behind them, and the city skyline beyond with windows beginning to light up. LIGHTING: Blue hour: deep cobalt sky with a thin band of orange at the horizon. CAMERA: Anamorphic cinema feel with a 40mm equivalent lens. FRAMING: vertical 9:16, faces and action in the center two thirds, headroom kept, nothing key in the bottom fifth. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Wide two shot. Daniel stands at the parapet looking at the skyline. Rosa steps out of the rooftop door behind him on the right, pauses, and holds up the two coffees. The camera slowly pushes in. Daniel turns his head toward her. SHOT 2, 3 to 6 seconds: Over Daniel's left shoulder onto Rosa in a medium close up. She hands him a cup, tilts her head, and speaks her line with a crooked, slightly nervous smile. Her face and mouth are clearly visible and lit by the string lights. SHOT 3, 6 to 10 seconds: Reverse over Rosa's right shoulder onto Daniel in a medium close up. He takes the cup in both hands, lets out a short laugh through his nose, and answers her, eyes wet but smiling. He then looks back at the skyline for the final beat. SOUND: Wind across the rooftop, distant traffic and a far away siren, a faint rumble of an L train, the creak of the rooftop door in shot 1, the soft crunch of gravel under Rosa's boots, cardboard cups brushing hands. A quiet, sparse piano note rises under the final second. DIALOGUE AND LIP SYNC: In SHOT 2 only Rosa speaks, softly and a little hesitant, mouth fully visible: "You still take it black?" In SHOT 3 only Daniel speaks, warm and slightly choked, mouth fully visible: "Some things never change, Rosa." Lip movements match each syllable precisely. The listener in each shot stays silent with a closed mouth, reacting only with eyes and breath. STYLE: Cinematic grade with teal shadows and warm highlights, fine film grain, gentle halation on the string lights, natural contrast. It should feel like a still from a modern indie drama. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

Prompt

ESPRESSO MACHINE PRODUCT FILM. A premium product film for a home espresso machine, focused on tactile detail, liquid physics, and satisfying sound design. One hand operates the machine; the product is the hero of every frame. THE SUBJECT: A compact brushed stainless steel espresso machine with a single group head, a walnut wood portafilter handle, a small round pressure gauge with a white dial, and a clean engraved word ORBIT on the front panel in thin capital letters. SETTING: A minimal kitchen counter made of pale grey concrete, a matte black tile backsplash, a small bag of coffee beans and a copper tamper at the side, a trailing plant softly out of focus in the background. Early morning, quiet and calm. LIGHTING: A single large soft source from the upper right like a window, with a black flag on the left for sculpted contrast. CAMERA: Macro and product lens look, 100mm macro for inserts and 50mm for the hero. Slow motorized slider moves, precise and smooth, no handheld shake. Motion is slow and deliberate, like a high end commercial. FRAMING: vertical 9:16, faces and action in the center two thirds, headroom kept, nothing key in the bottom fifth. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Macro close up. The hand locks the walnut handled portafilter into the group head with a firm quarter turn. A tiny puff of coffee dust drifts in the backlight. Camera slides slowly left to right. SHOT 2, 3 to 7 seconds: Extreme close up under the spout. Two thin streams of dark espresso begin to flow, turning honey brown as crema forms and swirls in the clear double wall glass. The pressure gauge needle rises in soft focus behind. Camera tilts down very slowly as the glass fills. SHOT 3, 7 to 10 seconds: Hero shot at 50mm. The full machine in three quarter view with the engraved ORBIT readable, steam curling from the group head, the finished espresso on the drip tray with a thick tiger striped crema. The hand lifts the cup out of frame on the final second. SOUND: A solid metal click and twist as the portafilter locks, the low hum and rhythmic pulse of the pump, a soft hiss of steam, the delicate trickle of espresso into glass, and a final porcelain free glass tap as the cup lifts. A very light warm ambient pad underneath. No voice. DIALOGUE AND LIP SYNC: There is no spoken dialogue in this clip. No one speaks and no narration is heard; the sound design carries it. STYLE: Premium commercial grade: warm neutral palette, deep espresso browns against brushed silver, crisp detail, subtle lens bloom on steam. Clean, modern, confident. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

Prompt

SOLAR EXPLAINER TALKING HEAD. A short educational talking head video in which an engineer explains one simple fact about home solar panels, with a cutaway to the panels. The priority is clear speech, accurate lip sync, and a trustworthy on-camera presence. THE SUBJECT: Priya, a South Asian American woman in her late thirties with shoulder length black hair in a low ponytail, rectangular tortoiseshell glasses, and a calm confident expression. SETTING: The backyard of a suburban home in Arizona on a clear day. LIGHTING: Bright late morning sun, softened by a large overhead diffusion so her face has even, flattering light with no squinting. The panels behind her show gentle sky reflections. Natural exposure with good detail in the bright sky. CAMERA: Documentary interview style. Shot 1 and shot 3: a locked off medium close up at eye level on a 50mm equivalent lens, subject slightly off center facing the lens. FRAMING: vertical 9:16, faces and action in the center two thirds, headroom kept, nothing key in the bottom fifth. SHOTS (total 10 seconds): SHOT 1, 0 to 4 seconds: Medium close up of Priya facing the camera directly, hard hat under her arm. She gives a small welcoming nod and says her first line, gesturing once toward the roof behind her with her free hand. SHOT 2, 4 to 7 seconds: Cutaway. The camera rises slowly above the roofline, revealing rows of solar panels glinting in the sun, a single white cloud reflected in the glass. Priya is not visible in this shot but her voice continues. SHOT 3, 7 to 10 seconds: Back to the same medium close up. Priya finishes her explanation looking straight into the lens, then smiles warmly and lifts her eyebrows as if to say that is all there is to it. SOUND: A light outdoor ambience: a soft breeze, distant birdsong, a faint neighborhood lawn mower far away. Her voice is close and clear like a lavalier microphone. A gentle, optimistic acoustic guitar bed at very low level. DIALOGUE AND LIP SYNC: Only Priya speaks, in a friendly, clear, teacherly tone at a relaxed pace. In SHOT 1, mouth fully visible facing the lens: "Solar panels make power from light, not heat." In SHOT 2 her voice continues off screen: "So cool, sunny days" In SHOT 3, mouth visible again, she finishes: "are actually their best days." Lip sync is exact in shots 1 and 3, every syllable matching. STYLE: Bright, clean, trustworthy look with natural colors, true blue sky, realistic skin tones, and crisp but not oversharpened detail. Modern educational channel aesthetic. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

Prompt

BIG SUR TRAVEL VLOG. A travel vlog moment on the California coast: a solo traveler reaches a cliffside viewpoint and shares her reaction, cut with sweeping views. It should feel spontaneous and joyful, with wind, sea, and a real voice. THE SUBJECT: Hannah, a white American woman in her mid twenties with strawberry blonde hair in a messy braid, light freckled skin, and a sun kissed nose. SETTING: A dirt trail on a grassy bluff in Big Sur, California, with the Pacific far below, white surf breaking on dark rocks, and the curve of a highway bridge in the distance. LIGHTING: Golden late afternoon sun low on the horizon to camera right, strong warm rim light on her hair, the ocean sparkling. The sky is soft blue fading to peach. CAMERA: Vlog style: shot 1 is a selfie from a small action camera on a short grip held at arm length, gently bouncing with her steps. Shot 2 is a smooth wide gimbal shot. Shot 3 returns to the selfie angle, now still. FRAMING: vertical 9:16, faces and action in the center two thirds, headroom kept, nothing key in the bottom fifth. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Selfie walking shot. Hannah hikes up the last part of the trail toward camera, slightly out of breath, glancing over her shoulder at the view, then back at the lens with wide excited eyes as she speaks. SHOT 2, 3 to 6 seconds: Wide shot from behind. She steps onto the edge of the viewpoint and the coastline opens up beneath her, surf rolling onto the rocks, the bridge in the distance. The camera glides forward and slightly up. SHOT 3, 6 to 10 seconds: Selfie again, now standing still with the ocean behind her. Wind blows loose strands of hair across her face. She brushes them away, laughs, and says her last line directly into the lens. SOUND: Strong coastal wind gusting across the action camera microphone, crunch of her boots on dirt, distant crashing surf, a seagull calling, her slightly breathless breathing in shot 1. In shot 2 the roar of the ocean swells. No music until a soft upbeat indie guitar fades in over the final two seconds. DIALOGUE AND LIP SYNC: Only Hannah speaks, in an excited, breathy, genuine voice. In SHOT 1, mouth visible to the selfie lens: "Okay, almost there." In SHOT 2 she is silent. In SHOT 3, mouth fully visible and close: "This is why I drove six hours." Lip movements match exactly, with a small laugh before the second line. STYLE: Warm, vivid travel grade: rich blues and golds, lifted shadows, a touch of action camera wide angle distortion on the selfie shots, natural skin. Authentic, not overproduced. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

Prompt

SMASH BURGER COOKING SHOW. A widescreen cooking segment in a diner kitchen where a cook smashes and flips a burger and talks to the camera. The focus is food physics, sizzling sound, and a charismatic line of dialogue with clean lip sync. THE SUBJECT: Marcus, a Black American man in his forties with a trimmed grey flecked beard, a shaved head, and a big easy grin. SETTING: A classic American diner kitchen with a wide black steel griddle, stainless steel shelves, ticket rail with paper orders, a red neon sign glowing softly on the back wall that reads OPEN in capital letters, and chrome fixtures. LIGHTING: Warm overhead kitchen lights with a stronger key from camera left, bright highlights on the steel. Rising smoke and steam are backlit so they glow. The neon adds a soft red accent in the background. CAMERA: Television cooking show look on a 35mm equivalent lens. Shot 1 is a medium wide from across the griddle. Shot 2 is a low macro angle at griddle level. Shot 3 is a medium close up of Marcus. Smooth dolly movement, crisp focus. FRAMING: widescreen 16:9, subjects on thirds, background carries depth. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Medium wide. Marcus drops a ball of beef on the hot griddle and immediately slams the press down with both hands. Smoke bursts up. He looks at the camera and starts talking. SHOT 2, 3 to 6 seconds: Low macro at griddle height. The patty edges turn lacy and crisp, fat bubbles and spits. The spatula scrapes under the patty and flips it, revealing a deep brown crust. A cheese slice lands on top and starts to melt down the sides. SHOT 3, 6 to 10 seconds: Medium close up. Marcus sets the cheesy patty on the toasted bun, points the spatula at the camera, and delivers his punchline with a grin, the red OPEN sign glowing behind him. SOUND: A loud aggressive sizzle when the beef hits the griddle and again under the press, the metallic clank of the press, the scrape of the spatula on steel, fat popping, the hum of an exhaust hood, and distant diner chatter and a bell from the front counter. DIALOGUE AND LIP SYNC: Only Marcus speaks, in a booming, playful, gravelly voice. In SHOT 1, facing the camera with his mouth clearly visible: "You gotta smash it hard." In SHOT 2 he is silent. In SHOT 3, mouth fully visible: "That crust is the whole burger." Lips sync exactly to each word, with a quick chuckle after the final line. STYLE: Rich, appetizing food grade: warm tones, deep browns, glossy melted cheese, crisp steel reflections. Sharp detail on the food with a slightly shallow background. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

Prompt

CLAYMATION FOX ANIMATION. A stop motion claymation short with a handmade look: a small fox baker in a cozy village bakery pulls a loaf from the oven and speaks one cheerful line. THE SUBJECT: Juniper, an anthropomorphic red fox made of matte modeling clay, about the proportions of a toddler, with a bright orange coat, a cream belly and muzzle, black clay paws, round glossy black bead eyes, and a fluffy white tipped tail. SETTING: A miniature clay and felt bakery set: a brick oven with a glowing orange mouth, wooden shelves with tiny clay loaves and croissants, a flour dusted wooden counter, a round window showing a painted cardboard village with a crooked church tower. LIGHTING: Warm tungsten practical light glowing out of the oven, a soft key from above as if a small studio lamp, gentle shadows with slightly hard edges like a real miniature set. A little flour dust hangs in the light beam. CAMERA: Miniature set photography with a shallow depth of field so the background softens like a tabletop. FRAMING: square 1:1, subject centered, nothing key near the corners. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Medium shot. Juniper opens the oven door with a wooden peel, the orange glow washing over her face, and slides out a golden round loaf. Her tail swishes behind her. SHOT 2, 3 to 6 seconds: Close up of the loaf on the counter as Juniper taps it with one paw. A puff of clay steam made of cotton wisps rises. Flour dust bounces off the counter. SHOT 3, 6 to 10 seconds: Close up of Juniper facing the camera, holding the loaf up proudly against her apron. She speaks her line, ears perking up, then gives a little bounce and a wink. SOUND: A creaky wooden oven door, a soft thump of the loaf on the counter, a hollow knock as she taps the bread, gentle crackle of the oven fire, a tiny bell over a door somewhere, and a whimsical plucked string and glockenspiel melody at low volume. DIALOGUE AND LIP SYNC: Only Juniper speaks, in a bright, sing song, high cartoon voice with a warm smile. In SHOT 3, facing camera with her clay mouth clearly visible: "Fresh bread for the whole village!" Her clay mouth changes shape in simple stop motion replacement style that matches each syllable. She is silent in shots 1 and 2. STYLE: Handmade claymation in the tradition of classic British stop motion studios: saturated warm oranges and blues, tactile surfaces, gentle imperfections, and a cozy storybook mood. Not glossy CGI, not photoreal. RULES: Realistic texture and physics unless the style says otherwise. Correct hands, five fingers, natural grip. Same faces, wardrobe, props and light in every shot. Clean cuts exactly at the timecodes. Only the named speaker moves their lips; others stay silent. No captions, subtitles, logos, watermarks or extra voices. Text on objects spelled exactly as written.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose FLUX 3 Video

Clips up to 20 seconds

Black Forest Labs built FLUX 3 Video to generate clips up to 20 seconds long. On Fliki you can set any length from 5 to 20 seconds, which leaves room for a full scene rather than a single moment.

Native audio in the same pass

The model generates dialogue, sound effects, and ambient sound together with the picture. Put spoken lines in quotes and describe the room tone, and the clip arrives with a matching soundtrack.

Multiple shots in one video

FLUX 3 Video can create several scenes and camera angles inside a single clip. Write timecoded shots and the model cuts between them while keeping the same people, wardrobe, and props.

720p or 1080p output

Render in HD for fast drafts or step up to 1080p for final delivery. Both tiers are available on Fliki in 16:9, 9:16, and 1:1.

One multimodal model

BFL trained FLUX 3 on image, video, and audio with a single set of weights. That shared training is why motion, lighting, and sound tend to agree with each other in the output.

First and last frame control

Start from an uploaded image, end on one, or both. Keyframes keep a product, a presenter, or a location exactly on model from the opening frame to the final one.

Follows shot-list direction

FLUX 3 Video handles dense direction: blocking, lens choice, light sources, and delivery notes. Prompts can run up to 3,000 characters, enough for a compact shot list with timings and dialogue.

Every social format

Compose natively for 16:9 YouTube and web video, 9:16 TikTok, Reels, and Shorts, or 1:1 feed posts from the same prompt.

How it works

How to generate a video with FLUX 3 Video

Getting a finished FLUX 3 Video clip with sound takes a few steps inside Fliki.

Fliki prompt input with a detailed shot description for FLUX 3 Video
Step 1

Write your prompt

Open Fliki and describe the video like a shot list: subject, setting, action, camera, lighting, and any dialogue in quotes. Prompts can be up to 3,000 characters, so keep each shot short and specific.

Fliki model selector with FLUX 3 Video chosen
Step 2

Select FLUX 3 Video as your model

Open the model selector and choose FLUX 3 Video. Fliki sends your prompt to the model with no extra setup.

Choose 16:9, 9:16, or 1:1 for FLUX 3 Video on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube and web, 9:16 for TikTok, Reels, and Shorts, or 1:1 for square feeds.

Set a 5 to 20 second duration for FLUX 3 Video on Fliki
Step 4

Set the duration

Pick any length from 5 to 20 seconds. Longer clips give room for several shots and a full line of dialogue; short clips iterate faster.

Upload first and last frame images for FLUX 3 Video on Fliki
Step 5

Add a first or last frame (optional)

Upload a starting image, an ending image, or both. FLUX 3 Video animates from the first frame and lands on the last, which keeps products and people on model.

Pick 720p or 1080p and generate with FLUX 3 Video on Fliki
Step 6

Select resolution and generate

Choose 720p or 1080p, then hit Generate. Preview the clip with its audio, download it, or drop it into a longer Fliki project.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

FLUX 3 Video FAQ

Frequently asked questions

Everything you need to know about generating with FLUX 3 Video inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
FLUX 3 Video · Free forever plan

Generate your next video with FLUX 3 Video.

Up to 20 seconds, several shots, and native audio from Black Forest Labs. Start free on Fliki, then upgrade when you are ready to render.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime