video model · by Google DeepMind

Gemini Omni Flash 1.1 AI Video Generator

Generate 3 to 10 second videos with synchronized audio using Gemini Omni Flash 1.1, Google's multimodal model for video generation. Direct the camera, write dialogue with lip-sync, set a first and last frame, and guide the look with reference images, all inside Fliki. Compare it side by side with all our AI video models before you render.

Generated with Gemini Omni Flash 1.1

A handful of Gemini Omni Flash 1.1 clips generated inside Fliki. No edits, no post.

Prompt

A 10 second vertical user-generated style testimonial for a vitamin C face serum, filmed as if on a phone by the creator herself in her own bathroom. It should feel honest, warm, and unpolished in the way real creator videos are, while the product and the speech stay crisp and perfectly in sync. Three shots, with dialogue delivered straight to camera. THE SUBJECT: Maya, a Black American woman around 29, with shoulder-length natural coils pulled back by a soft pale yellow headband, clear medium-deep brown skin with a healthy morning glow, small gold hoop earrings, and a loose white ribbed tank top. Her nails are short with a nude polish. She holds a 30 ml amber glass dropper bottle with a matte white cap and a plain white label that reads GLOW C in simple black capitals. The bottle, the label, the headband, and the earrings stay exactly the same in every shot. SETTING: A small, tidy apartment bathroom in the morning. White subway tile, a round mirror with a thin brass frame behind her, a potted pothos on the counter to one side, a folded sage green hand towel, and a ceramic cup holding a toothbrush. Nothing branded is visible except the GLOW C label. LIGHTING: Soft daylight from a frosted window on camera left, slightly warm, wrapping gently across her face. A faint natural highlight on her cheekbones and forehead. No ring light reflections in her eyes, no harsh shadows, no color grading beyond a clean, true-to-life look. CAMERA: Handheld phone feel with very small natural sway, as if the creator holds the phone at arm's length or has propped it on the counter. Focus is always sharp on her face or on the bottle when it is featured. Shallow but realistic phone depth of field. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3.5 seconds: Medium close-up, phone propped at eye level. Maya leans into frame from the right, smiles, and lifts the GLOW C bottle beside her cheek with the label facing the lens. She speaks her first line directly to camera. SHOT 2, 3.5 to 6.5 seconds: Close-up insert. Her left hand squeezes the dropper and places three drops onto her fingertips of the right hand; the serum is a clear, slightly golden liquid that catches the window light. She pats it gently onto her cheek. The label stays readable in soft focus on the counter. SHOT 3, 6.5 to 10 seconds: Back to the medium close-up from Shot 1. She turns her face slightly to show the dewy finish on her cheek, then looks back into the lens, delivers her final line with a small laugh, and gives a tiny nod as the clip ends. DIALOGUE AND LIP-SYNC: Maya is the only speaker. Her mouth is fully visible and unobstructed whenever she speaks, and her lip movement matches every syllable exactly. In Shot 1 she says, in a friendly, conversational tone: "Okay, this is the serum I keep repeating." In Shot 2 she is silent. In Shot 3 she says, relaxed and a little amused: "Two weeks in, and look at this glow." Natural American accent, normal speaking speed, no exaggerated enunciation. SOUND: Close, intimate phone-mic acoustics of a small tiled room with a slight natural reverb. Soft drip of the dropper and a faint tap of fingertips on skin in Shot 2. A quiet room tone underneath. No music. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The bottle label reads GLOW C exactly, in the same font and placement every time it appears. - Skin texture looks real, with visible pores; do not airbrush or blur her skin. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical cinematic short about two brothers meeting again on a city rooftop at dusk after years apart. It should play like the emotional middle of a feature film: restrained performances, precise camera work, and one short exchange of dialogue. Three shots. THE SUBJECT: Daniel, a Mexican American man around 35, with short black hair, a neatly trimmed beard, a charcoal wool overcoat over a navy crewneck sweater. Luis, his younger brother around 30, with a buzz cut, a light stubble, a worn tan canvas work jacket over a gray hoodie. Daniel holds a paper cup of coffee with a plain brown sleeve. Their wardrobe and hair stay exactly the same in every shot. SETTING: A flat gravel rooftop in Chicago at dusk. A low brick parapet, a rusted water tower on the neighboring building, and the downtown skyline behind them with lights just starting to come on. A string of unlit bulbs sags between two poles. A light wind moves their hair and coat hems. LIGHTING: Blue hour. Cool ambient sky light fills the scene while the warm orange of the city windows glows behind them. A faint warm practical light from a rooftop door spills across their faces from camera right. Gentle contrast, rich blacks, film-like color. CAMERA: Cinematic, stable, and deliberate. Shot 1 uses a slow dolly-in. Shot 2 is an over-the-shoulder hold. Shot 3 is a slow orbital move around the two of them. Anamorphic feel with soft background bokeh on the skyline lights, subtle film grain. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3.5 seconds: Wide vertical shot from behind Daniel as he steps out of the rooftop door. The camera dollies in slowly as Luis, standing at the parapet, turns around to face him. Both stop, a few steps apart. SHOT 2, 3.5 to 6.5 seconds: Over Daniel's shoulder onto Luis in medium close-up. Luis swallows, then speaks his line quietly, eyes glistening but not crying. SHOT 3, 6.5 to 10 seconds: Medium two-shot, the camera slowly orbiting a quarter turn to the left. Daniel answers, holds out the coffee cup; Luis takes it, and the two share a short, awkward laugh that turns into a brief one-armed hug as the skyline lights continue to flicker on. DIALOGUE AND LIP-SYNC: Luis speaks first, then Daniel. In Shot 2, Luis, soft and hesitant, with his mouth clearly visible: "You actually came." In Shot 3, Daniel, warm and steady, facing three quarters toward camera with his mouth visible: "Of course I came. I brought coffee." Lip movement matches each word precisely. American accents, natural pauses, no overacting. SOUND: Open rooftop air: a light steady wind, distant city traffic, a faraway siren fading, the gravel crunching under Daniel's shoes in Shot 1, the rooftop door clicking shut behind him. A low, sparse piano note sustains softly under Shot 3 only. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The coffee cup stays in Daniel's hand until he hands it over in Shot 3, and then stays in Luis's hand. - Performances are subtle and grounded; no melodramatic crying or shouting. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical product film for a pair of matte black wireless earbuds and their charging case. Pure product storytelling with no people except one hand. It should feel like a premium tech launch ad: controlled light, precise camera moves, satisfying mechanical sound design. Three shots with a snap-zoom and an orbital move. THE SUBJECT: A pebble-shaped charging case in matte black with a soft-touch finish, a thin brushed aluminum hinge line, and a tiny white status light on the front. Inside, two matte black earbuds with a small brushed aluminum ring on each stem. The case lid is engraved with the word PULSE in small, clean capital letters. A single hand, belonging to a light-skinned person with short clean nails and no jewelry, appears only in Shot 2. SETTING: A seamless dark graphite studio surface that curves up into an infinite background. A shallow layer of fine mist drifts low across the surface. No other objects. LIGHTING: Low-key studio lighting. A long narrow softbox from above creates a clean gradient highlight along the curve of the case. A thin cool-white rim light from behind separates the product from the background. A tiny warm kicker catches the aluminum ring. Deep, clean blacks, no dust, no fingerprints. CAMERA: Motion-control precision. Shot 1 is a slow push-in. Shot 2 includes a fast snap-zoom onto the earbud. Shot 3 is a smooth 180 degree orbit around the open case. Macro lens character with a very shallow depth of field and creamy falloff. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: The closed case rests on the surface in the mist. The camera pushes in slowly and the white status light pulses once. The engraved word PULSE comes into sharp focus on the lid. SHOT 2, 3 to 6 seconds: The hand enters from the top of frame and flips the lid open with the thumb; the lid swings open with a smooth damped motion. The camera snap-zooms onto the left earbud as it lifts slightly out of its magnetic seat, the aluminum ring catching a streak of light. SHOT 3, 6 to 10 seconds: The hand has left the frame. The open case sits alone as the camera orbits 180 degrees around it at product height, the rim light sliding across the lid and both earbuds. The move eases out and settles on a clean three-quarter hero angle as the mist thins. DIALOGUE AND LIP-SYNC: There is no dialogue in this video. SOUND: Tight, premium product sound design. A soft low hum under Shot 1 with a gentle electronic chime synced exactly to the status light pulse. A crisp magnetic click as the lid opens and a subtle whoosh on the snap-zoom in Shot 2. A slow airy swell during the orbit in Shot 3, ending on one clean bass tone as the move settles. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The engraving reads PULSE exactly, and the case geometry, finish, and hinge are identical in every shot. - Surfaces are flawless and physically plausible; reflections move correctly with the camera. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical talking-head explainer for a personal finance channel. A presenter explains one simple idea about credit scores directly to camera. It should feel like a confident, trustworthy creator video: clean framing, clear speech, one cutaway to a whiteboard. Three shots. THE SUBJECT: Priya, an Indian American woman around 34, with long dark brown hair worn straight over one shoulder, light makeup, thin rectangular tortoiseshell glasses, a mustard yellow knit cardigan over a white blouse. She holds a black whiteboard marker in her right hand in all three shots. SETTING: A bright home office. Behind her, a white wall with a small framed abstract print, a wooden bookshelf with a few plants and books, and to her left a freestanding whiteboard. On the whiteboard, written neatly in blue marker, are the words PAY ON TIME with an underline, and below it a simple hand-drawn arrow pointing up. That is the only text on the board. LIGHTING: Clean and flattering key light from a large window on camera right, balanced with a soft white fill on the left. Warm, natural skin tones. The whiteboard is evenly lit with no glare. CAMERA: Locked-off tripod shots with a slight slow push-in during Shot 1 and Shot 3. Sharp focus on her eyes. A realistic 35 mm equivalent lens with the background gently out of focus. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3.5 seconds: Medium close-up of Priya centered in frame, looking into the lens. She raises the marker slightly like a pointer and delivers her first line with a quick, confident smile. SHOT 2, 3.5 to 6.5 seconds: Medium shot, angled to include the whiteboard beside her. She taps the words PAY ON TIME twice with the capped marker and then traces the upward arrow while speaking. SHOT 3, 6.5 to 10 seconds: Back to the medium close-up. She turns back to camera, lowers the marker, and finishes her line with a small nod and a raised eyebrow. DIALOGUE AND LIP-SYNC: Priya is the only speaker. Her mouth is fully visible in every shot and the lip-sync is exact. Shot 1, friendly and direct: "Want a better credit score fast?" Shot 2, clear and explanatory, as she taps the board: "Payment history is the biggest factor." Shot 3, warm and reassuring: "So set up autopay today." Natural American accent with a light, confident rhythm. SOUND: Quiet treated room acoustics, close lavalier-mic quality voice. Two soft plastic taps on the whiteboard in Shot 2 and a faint squeak as she traces the arrow. Very light room tone. No music. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The whiteboard text reads PAY ON TIME exactly, and it does not change between shots. - Her glasses stay on and keep the same position; reflections in the lenses stay subtle. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical travel vlog moment in Kyoto on a quiet autumn morning. A traveler walks, reacts, and speaks one short line to the camera. It should feel like a premium travel creator reel: golden light, real places, natural handheld energy. Three shots including an orbital move. THE SUBJECT: Jordan, a white American man around 27, with wavy light brown hair, a short beard, a rust orange puffer vest over a cream henley, dark olive chinos, and a small black canvas camera sling bag across his chest. He wears the same outfit and bag in every shot. SETTING: The stone-paved lanes of Higashiyama in Kyoto, with traditional two-story wooden machiya houses, dark tiled roofs, paper lanterns hanging unlit by the doorways, and red and orange maple trees in full autumn color. The five-story Yasaka Pagoda rises at the top of the lane. A few distant pedestrians, no crowds. LIGHTING: Early morning sun low in the sky from behind the pagoda, creating warm golden backlight, long soft shadows on the stone, and glowing edges on the maple leaves. The shadow side of the lane is cool and blue for natural contrast. CAMERA: Shot 1 is a smooth gimbal follow from behind. Shot 2 is a slow orbital move around Jordan. Shot 3 is a selfie-style handheld angle held by Jordan himself at arm's length, slight natural shake. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3.5 seconds: The camera follows Jordan from behind as he walks up the gently sloping lane toward the pagoda, leaves drifting down around him. He slows as the pagoda comes fully into view. SHOT 2, 3.5 to 6.5 seconds: Medium shot as the camera orbits a half circle around Jordan, revealing his face in golden backlight as he looks up at the pagoda with a quiet smile and exhales in awe. SHOT 3, 6.5 to 10 seconds: Selfie-style close-up, the pagoda framed over his shoulder. He speaks his line directly into the lens, then turns the camera slightly to include more of the pagoda as the clip ends. DIALOGUE AND LIP-SYNC: Jordan is the only speaker, and he speaks only in Shot 3. His mouth is clearly visible and the lip-sync is exact. He says, softly and a little breathless, as if not wanting to disturb the quiet: "Six a.m. in Kyoto. Totally worth it." Natural American accent, relaxed pace. SOUND: Quiet morning ambience: footsteps on stone, a gentle breeze moving the maple leaves, a distant temple bell ringing once during Shot 2, a few birds. In Shot 3 his voice is close to the mic with natural outdoor acoustics. No music. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The Yasaka Pagoda keeps the same shape, five roof tiers, and position relative to the lane in every shot. - Leaves fall with natural physics and do not pass through people or objects. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second widescreen food film set in a classic American diner at breakfast rush. A short-order cook makes a smash burger with crisp edges and melted cheese, and talks to a regular across the counter. It should feel like a mouthwatering documentary food moment with real sizzle and grease. Three shots. THE SUBJECT: Earl, a Black American man around 58, with close-cropped gray hair and a gray mustache, a white short-sleeve cook's shirt with rolled sleeves, a striped blue and white apron, and a white paper diner cap. He holds a heavy stainless steel smash press and a flat metal spatula. His outfit and tools stay the same in every shot. SETTING: A 1950s style diner kitchen line visible over a chrome and red vinyl counter. A seasoned black flat-top griddle, a steel order rail with paper tickets, stacked white ceramic plates, a squeeze bottle of yellow mustard. A red neon sign on the back wall reads OPEN in cursive; that is the only readable text. LIGHTING: Warm overhead diner fluorescents mixed with the red glow of the neon sign behind. Steam and smoke rise from the griddle and catch a shaft of morning sunlight from the front window on camera left, making the smoke visible and golden. CAMERA: Shot 1 is a macro close-up at griddle level. Shot 2 is a medium shot across the counter with a slow dolly left. Shot 3 is a slow push-in to a hero close-up of the finished burger. Shallow depth of field with rich, appetizing color. FRAMING (LANDSCAPE 16:9): Compose for a widescreen display. Use the width deliberately: place the subject on a rule-of-thirds line with the direction of movement or gaze leading into open space. Keep horizons level unless a shot calls for a tilt. Leave a comfortable margin on all sides so nothing important touches the frame edge. Do not letterbox or add black bars. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Macro close-up at griddle height. A ball of ground beef hits the flat-top and Earl slams the press down, flattening it into a thin patty; fat spits and the edges immediately lace and brown. SHOT 2, 3 to 6.5 seconds: Medium shot across the counter. Earl flips the patty with the spatula, lays a slice of American cheese on top, and glances up to talk to the unseen regular off camera right while the cheese starts to melt. SHOT 3, 6.5 to 10 seconds: Slow push-in as he sets the finished burger on a toasted potato bun on a white plate and slides it forward across the counter toward the camera. Melted cheese drapes over the crispy edges; a wisp of steam rises. DIALOGUE AND LIP-SYNC: Earl is the only speaker, and he speaks in Shot 2 and Shot 3. His mouth is clearly visible and lip-sync is exact. Shot 2, easygoing and proud, eyes toward the regular: "Crispy edges, just how you like it." Shot 3, with a grin as he slides the plate: "Order up, my friend." A warm, gravelly Southern American voice. SOUND: Loud authentic diner audio: a violent sizzle when the beef hits the griddle and the press slams, the scrape of the spatula on steel, the soft hiss of melting cheese, a ticket bell ringing once, low background chatter and clinking coffee mugs. The ceramic plate sliding on the counter in Shot 3. No music. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The neon sign reads OPEN exactly and stays in the same place. - Food looks real and appetizing with true browning and grease; no plastic sheen. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second widescreen sports moment: the final meters of a high school 400 meter race under stadium lights, followed by the winner's reaction and a line to her coach. It should feel like a sports broadcast and a documentary in one, with real speed, sweat, and emotion. Three shots including a dolly-zoom. THE SUBJECT: Aaliyah, a Black American teen sprinter around 17, with her hair in a high braided bun, a sleeveless maroon racing singlet with a white bib reading 214 pinned to the front, black compression shorts, and neon green spikes. Her coach, Coach Ramirez, a Latino man around 45 with a gray baseball cap, a black windbreaker, and a stopwatch on a lanyard. Wardrobe, bib number, and hair stay identical in every shot. SETTING: A high school outdoor track at night with a red rubber surface, white lane lines, and a finish line with a small digital timer. Metal bleachers with a scattering of cheering parents. Tall stadium light towers glow against a deep navy sky. LIGHTING: Hard white stadium light from high above creating crisp shadows and a bright sheen of sweat on skin. The background is darker, with the light towers flaring slightly. Night air has a faint haze that catches the light beams. CAMERA: Shot 1 is a fast tracking shot alongside her on a camera car. Shot 2 is a dolly-zoom on her face just after she crosses the line, the background stretching away while she stays the same size. Shot 3 is a handheld medium two-shot. FRAMING (LANDSCAPE 16:9): Compose for a widescreen display. Use the width deliberately: place the subject on a rule-of-thirds line with the direction of movement or gaze leading into open space. Keep horizons level unless a shot calls for a tilt. Leave a comfortable margin on all sides so nothing important touches the frame edge. Do not letterbox or add black bars. SHOTS (total 10 seconds): SHOT 1, 0 to 3.5 seconds: The camera tracks alongside Aaliyah in lane 4 at full sprint in the last 30 meters, arms pumping, jaw set. Two other runners in blue singlets are a stride behind her. She leans through the finish line. SHOT 2, 3.5 to 6 seconds: Dolly-zoom on Aaliyah's face as she slows past the line and looks up at the digital timer, eyes wide in disbelief, chest heaving. SHOT 3, 6 to 10 seconds: Handheld medium two-shot as Coach Ramirez jogs up holding the stopwatch. She grabs his arm and speaks first; he answers, laughing, and taps the stopwatch face. DIALOGUE AND LIP-SYNC: Aaliyah speaks first, then Coach Ramirez, both only in Shot 3, with mouths clearly visible and exact lip-sync. Aaliyah, out of breath, excited, voice cracking slightly: "Coach, did I break it?" Coach Ramirez, grinning, warm and loud over the crowd: "By half a second. You did it." Natural American accents. SOUND: Fast rhythmic spike footfalls and heavy breathing in Shot 1, a rising crowd roar as she crosses the line, a short air horn blast. In Shot 2, the crowd briefly drops away into a muffled heartbeat and her breathing, then rushes back in at the start of Shot 3. The coach's lanyard jingles. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound is tied to a visible cause and lands on the exact frame of the action that makes it, such as footsteps on foot contact, clicks on contact, and sizzles on impact. Room acoustics match the space shown in each shot, and the loudness of a sound follows its distance from the camera. When someone speaks, only that person's lips move, the jaw and cheeks move naturally with the words, and breaths happen between phrases. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The bib always reads 214. The other runners wear blue singlets and never merge with her. - Running mechanics are realistic with correct foot strikes and arm swings at true speed. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose Gemini Omni Flash 1.1

Native audio in the same pass

Every clip comes with a generated soundtrack: ambience, sound effects, and speech that line up with what happens on screen. Describe the sound you want in the prompt and skip the separate audio step.

Dialogue with lip-sync

Write lines in quotes and say who speaks them. Google built Gemini Omni Flash 1.1 for character dialogue and lip-sync, so talking-head explainers and short scenes play back with matching mouth movement.

First and last frame control

Set the opening frame, or both the opening and closing frames, and the model generates the motion in between. Useful for product reveals, transitions, and camera moves that must land on an exact composition.

Reference images for consistency

Add reference images of a person, product, or look and the model carries them into the shot. It keeps a character or packshot recognizable from one clip to the next.

Precise camera direction

Google highlights camera moves such as dolly zooms, orbital rotations, and snap zooms. Name the move, its speed, and where it ends, and the model follows it through the clip.

Follows long, detailed prompts

Gemini Omni Flash 1.1 handles rich prompts that spell out subject, wardrobe, lighting, timing, and sound. Fliki accepts prompts up to 40,000 characters for this model, so a full shot list fits.

720p and 1080p output

Render at 720p for fast drafts or 1080p for final delivery. Output is 24 fps, a natural film cadence for social, web, and ads.

Vertical and landscape framing

Generate in 9:16 for TikTok, Reels, and Shorts or 16:9 for YouTube and web. Each ratio is composed natively, so the subject stays properly framed.

How it works

How to generate a video with Gemini Omni Flash 1.1

Getting a production-ready video out of Gemini Omni Flash 1.1 takes under a minute inside Fliki. Follow these six steps.

Fliki prompt input with a detailed text-to-video description for Gemini Omni Flash 1.1
Step 1

Write your prompt

Open Fliki and describe the scene in plain language. Name the subject, setting, camera move, lighting, and any sound or dialogue. Gemini Omni Flash 1.1 follows long, detailed prompts, so write it like a shot list.

Fliki model selector with Gemini Omni Flash 1.1 chosen for AI video generation
Step 2

Select Gemini Omni Flash 1.1 as your model

Open the model selector and choose Gemini Omni Flash 1.1. Fliki sends your prompt straight to Google's model with no API keys or extra setup.

Choose 16:9 or 9:16 aspect ratio for Gemini Omni Flash 1.1 video generation on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube and landscape web, or 9:16 for TikTok, Reels, and Shorts. These are the two ratios Gemini Omni Flash 1.1 supports.

Set a 3 to 10 second duration for Gemini Omni Flash 1.1 on Fliki
Step 4

Set the duration

Pick a clip length from 3 to 10 seconds. Short clips are quick to iterate on; 10 seconds gives the model room for a full beat with camera movement and a line of dialogue.

Upload first and last frame images or reference images for Gemini Omni Flash 1.1 on Fliki
Step 5

Add frames or references (optional)

Upload a first frame, or a first and last frame, to control where the shot starts and ends. You can also add reference images to keep a character, product, or style consistent.

Pick 720p or 1080p and generate a video with Gemini Omni Flash 1.1 on Fliki
Step 6

Select resolution and generate

Pick 720p for faster turnaround or 1080p for sharper detail, then hit Generate. Your clip arrives with audio, ready to preview, download, or drop into a longer Fliki project.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

Gemini Omni Flash 1.1 FAQ

Frequently asked questions

Everything you need to know about generating with Gemini Omni Flash 1.1 inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
Gemini Omni Flash 1.1 · Free forever plan

Generate your next video with Gemini Omni Flash 1.1.

Videos with native audio, lip-synced dialogue, and precise camera direction from Google's multimodal video model. Free to start, no credit card required.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime