video model · by Lightricks

LTX-2.5 Fast AI Video Generator

Generate videos up to 20 seconds with synchronized audio using LTX-2.5 Fast, the speed tier of Lightricks' open-weights video model. Multi-shot scenes, first and last frame control, and 720p or 1080p output, all inside Fliki. Compare it side by side with all our AI video models before you render.

Generated with LTX-2.5 Fast

A handful of LTX-2.5 Fast clips generated inside Fliki. No edits, no post.

Prompt

A 10 second vertical user-generated style ad for a canned oat milk cold brew, filmed as a quick creator video in a sunny apartment kitchen. Four fast connected shots in one generation, the same creator, can, and kitchen throughout, with one short spoken line and crisp product sound. THE SUBJECT: Sofia, a Puerto Rican American woman around 26, with long dark curly hair in a loose claw clip, a few curls framing her face, warm olive skin, a light lavender oversized crewneck sweatshirt with the sleeves pushed up, and thin silver rings on two fingers. She holds a slim 12 ounce aluminum can in pale sky blue with a white wave graphic and the word DRIFT printed vertically in bold white capitals. The can, the sweatshirt, the clip, and the rings stay identical in every shot. SETTING: A small bright apartment kitchen with white cabinets, a light oak butcher block counter, a trailing pothos on the windowsill, a glass jar of oats, and a tall clear glass filled with ice cubes. The window behind her shows soft green trees. No other brands visible. LIGHTING: Late morning sunlight through the window on camera right, warm and bright, casting soft window-frame shadows on the counter. Condensation on the can glitters in the light. Natural, clean color with gentle contrast. CAMERA: Handheld phone creator energy with small natural movement, quick purposeful cuts, sharp focus on the can whenever it is featured and on her face when she speaks. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 2 seconds: Close-up. Sofia's hand pulls the cold, sweating DRIFT can out of the fridge door, droplets running down the label; the fridge light spills onto the can. SHOT 2, 2 to 4.5 seconds: Macro close-up. Her thumb cracks the tab open; a tiny mist puffs from the opening. SHOT 3, 4.5 to 7 seconds: Medium close-up from counter height. She pours the cold brew over the ice in the tall glass; creamy brown coffee swirls around the cubes in slow ribbons. SHOT 4, 7 to 10 seconds: Medium close-up of Sofia centered in frame. She takes a sip, closes her eyes for a beat, then looks into the lens, holds the can up next to her face with the DRIFT label facing camera, and says her line with a small grin. DIALOGUE AND LIP-SYNC: Sofia is the only speaker, and she speaks only in Shot 4. Her mouth is clearly visible and her lips match every word. She says, bright and casual: "Honestly? Better than my coffee shop." Natural American accent, conversational pace. SOUND: Kitchen ambience: the fridge door suction pop and hum in Shot 1, a sharp satisfying can crack and fizz in Shot 2, liquid pouring over ice with clinking cubes in Shot 3, a small appreciative hum after her sip in Shot 4. No music. IMAGE QUALITY AND TEXTURE: Real camera footage, not a render. True-to-life color and skin tones, believable fabric and hair, natural motion blur, real lens depth of field, and stable backgrounds that never shimmer or change between shots. AUDIO TO PICTURE SYNC: Every sound lands on the frame of the action that causes it, room acoustics match each space, and only the speaker's lips move when a line is spoken. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The can reads DRIFT exactly, vertical, in white capitals, sharp and legible whenever the label faces camera. - Liquid pours and ice behave with real physics; condensation drips downward. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical cinematic action sequence: a courier sprints through a crowded, rain-soaked night market with a stolen-looking package, with a pursuer close behind. Built to show dense detail holding together in fast motion and crowds. Four connected shots with the same characters and location. THE SUBJECT: Kenji, a Japanese American man around 30, with a short undercut and wet black hair, a black waterproof shell jacket with reflective silver piping on the shoulders, a gray hoodie beneath, and a yellow insulated delivery backpack. He clutches a small brown paper package tied with red string. The pursuer is a tall man in a long dark trench coat and a flat cap whose face stays mostly in shadow. Wardrobe and the package stay identical across shots. SETTING: A narrow night market alley in a rainy American city Chinatown. Red paper lanterns strung overhead, steam rising from food stalls, stacked crates of produce, puddles reflecting neon. A vertical neon sign on the left wall reads NOODLES in pink capitals. A crowd of shoppers with umbrellas fills the lane. LIGHTING: Wet neon night. Magenta, cyan, and warm lantern light reflecting in puddles and on wet jackets. Rain streaks catch the light as bright lines. Deep shadows between stalls. CAMERA: Kinetic but readable. Shot 1 is a low tracking shot in front of Kenji moving backward. Shot 2 is a whip pan. Shot 3 is an overhead top-down angle. Shot 4 is a close-up handheld hold. Shutter gives natural motion blur on rain and limbs. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Low tracking shot in front of Kenji as he sprints straight toward camera through the crowd, ducking under an umbrella and splashing through a puddle, the package tight against his chest. SHOT 2, 3 to 5 seconds: Whip pan to the pursuer shouldering past shoppers at the far end of the lane, then back to Kenji vaulting over a stack of crates. SHOT 3, 5 to 7.5 seconds: Overhead shot looking straight down on the lane: a sea of colored umbrellas parts as Kenji weaves through, the yellow backpack easy to track. SHOT 4, 7.5 to 10 seconds: Close-up of Kenji pressed into a dark doorway under the NOODLES sign, breathing hard, rain dripping off his hood. He glances back, then looks down at the package and speaks under his breath. DIALOGUE AND LIP-SYNC: Kenji is the only speaker, in Shot 4. His mouth is visible in the neon light and lip-sync is exact. He whispers, tense: "Not tonight. Not this one." American accent. SOUND: Heavy rain on awnings and umbrellas, splashing footsteps, crowd murmur and a few startled shouts, sizzling woks from the stalls, crates rattling as he vaults them. In Shot 4 the ambience drops to rain and his ragged breathing. A low pulsing synth bass drives Shots 1 to 3. IMAGE QUALITY AND TEXTURE: Real camera footage, not a render. True-to-life color and skin tones, believable fabric and hair, natural motion blur, real lens depth of field, and stable backgrounds that never shimmer or change between shots. AUDIO TO PICTURE SYNC: Every sound lands on the frame of the action that causes it, room acoustics match each space, and only the speaker's lips move when a line is spoken. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The neon sign reads NOODLES exactly. - Crowd members keep distinct faces, bodies, and umbrellas; no people merging or duplicating. - Fast motion stays sharp enough to read; limbs never smear into each other. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical product film for a lightweight running shoe, mixing studio product shots with one real run on a city bridge. Four connected shots in a single generation with the same shoe, colors, and branding throughout. THE SUBJECT: A running shoe with a breathable coral orange knit upper, a thick white foam midsole with a subtle wavy sidewall texture, a black rubber outsole, and white flat laces. On the heel tab, the word STRIDE is printed in small black capitals. In Shot 4 it is worn by a runner: Marcus, a Black American man around 32, with short twists, a black running tank, black shorts, and a white sports watch. SETTING: Shots 1 to 3: a seamless warm sand colored studio sweep with a shallow puddle of water on the floor for splash. Shot 4: the Brooklyn Bridge pedestrian walkway at sunrise, wooden planks and stone towers, Manhattan skyline in soft haze. LIGHTING: Studio: a large soft key from the top left, a warm rim light from behind, clean shadow under the shoe. Bridge: low golden sunrise from the side, long shadows across the planks, warm glow on the knit. CAMERA: Shot 1 is a slow orbit. Shot 2 is a macro slider move. Shot 3 is a locked high-speed feel with a real-time splash. Shot 4 is a low tracking shot at ankle height moving with the runner. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 2.5 seconds: The single shoe floats a few centimeters above the studio floor, slowly rotating as the camera orbits the opposite way, showing the coral knit and the white midsole. SHOT 2, 2.5 to 5 seconds: Macro slide along the side of the shoe across the knit weave and the wavy foam texture, ending on the heel tab where STRIDE is sharp and legible. SHOT 3, 5 to 7 seconds: The shoe drops onto the shallow puddle; water splashes outward in a crown of droplets that catch the rim light, the foam compressing slightly and rebounding. SHOT 4, 7 to 10 seconds: Ankle-height tracking shot beside Marcus as he runs across the bridge planks at an easy tempo in the same shoes, each stride landing crisply, the skyline glowing behind. DIALOGUE AND LIP-SYNC: There is no dialogue in this video. SOUND: Clean studio air and a soft whoosh during the orbit, a tactile brush sound as the camera slides across the knit, a bright water splash and a soft foam thud in Shot 3. In Shot 4, rhythmic footfalls on wooden planks, light breathing, distant city traffic and seagulls. A light, upbeat electronic pulse builds across all four shots. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - STRIDE is spelled exactly and appears only on the heel tab. - The shoe colors and design are identical in studio and on the runner. - Water splashes follow real physics. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical fitness explainer. A coach gives one quick form tip for bodyweight squats, speaking to camera and then demonstrating. Clear, bright, and trustworthy. Three connected shots with the same coach and room. THE SUBJECT: Coach Tasha, a white American woman around 38, with a blonde high ponytail, athletic build, a teal sports bra under an open light gray zip hoodie, black high-waisted leggings, and white trainers. A small black fitness tracker on her left wrist. Wardrobe and hair stay identical across shots. SETTING: A bright home gym in a converted garage: light gray rubber floor tiles, a rack of colorful dumbbells, a folded yoga mat, a large wall mirror, and a plant in the corner. A small whiteboard on the wall reads KNEES OVER TOES in green marker; that is the only text. LIGHTING: Clean diffused daylight from a roll-up garage door on camera left plus soft overhead LED panels. Even, flattering light with no harsh shadows. Natural color. CAMERA: Tripod-mounted with a gentle push-in in Shot 1, a side-profile locked shot in Shot 2, and a low front angle in Shot 3. Sharp focus on her face and knees. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3.5 seconds: Medium shot of Tasha centered, facing camera, hands on hips. She smiles and speaks her first line as the camera pushes in slightly. SHOT 2, 3.5 to 7 seconds: Side profile full-body shot as she performs one slow bodyweight squat: hips back, chest up, knees tracking over her toes, then stands back up. The whiteboard is visible behind her. SHOT 3, 7 to 10 seconds: Low front angle medium shot. She finishes standing, points to her knees, and delivers her last line with an encouraging nod. DIALOGUE AND LIP-SYNC: Tasha is the only speaker. Mouth clearly visible with exact lip-sync. Shot 1, upbeat and clear: "Squats hurting your knees? Try this." Shot 3, encouraging: "Hips back first, knees follow." Natural American accent, coach-like energy without shouting. SOUND: Garage acoustics with a light echo, her voice close and clear. Soft sneaker squeaks on the rubber tiles and a controlled exhale during the squat. Faint birdsong from the open door. No music. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound lands on the frame of the action that causes it, room acoustics match each space, and only the speaker's lips move when a line is spoken. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The whiteboard reads KNEES OVER TOES exactly. - Squat form is anatomically correct and demonstrates good technique; joints bend naturally. - Her ponytail swings naturally with movement and settles when she stops; the dumbbell rack and mirror stay in the same place behind her in every shot. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second vertical travel reel of a couple at the Grand Canyon South Rim at sunrise, from arrival to a shared moment on the edge. Four connected shots with consistent people, wardrobe, and light, showing the scale and dense detail of the landscape. THE SUBJECT: Ahmed, a Pakistani American man around 31, with short black hair, a trimmed beard, a forest green fleece jacket, tan hiking pants, and a gray beanie. His wife Leah, a white American woman around 30, with a long auburn braid, a mustard yellow puffer jacket, black leggings, and brown hiking boots. Both carry small daypacks. Wardrobe stays identical across shots. SETTING: Mather Point on the Grand Canyon South Rim at dawn. Layered red, orange, and violet canyon walls stretching to the horizon, a thin ribbon of the Colorado River far below, a low metal guardrail, scattered juniper trees, and frost on the rocks. LIGHTING: Sunrise breaking over the eastern rim: the first warm rays skim across the canyon, lighting ridge tops in bright orange while the depths stay in cool blue shadow. Long soft shadows from the couple across the rock. CAMERA: Shot 1 is a wide drone shot pushing over the rim. Shot 2 is a gimbal follow. Shot 3 is a slow crane up. Shot 4 is a close handheld two-shot. FRAMING (VERTICAL 9:16): Compose for a phone screen held upright. Keep the main subject centered horizontally in every shot, with the eyes or the key product detail sitting in the upper third of the frame. Leave safe margins of roughly ten percent at the top and fifteen percent at the bottom so nothing important falls under app buttons, captions, or the progress bar. Avoid wide empty sky or floor; use the height of the frame for the body, the gesture, and the object in hand. Do not letterbox, do not pillarbox, and do not crop a landscape composition into vertical. SHOTS (total 10 seconds): SHOT 1, 0 to 3 seconds: Drone shot low over the juniper trees pushing out past the rim to reveal the vast canyon glowing in the first light, the two tiny figures walking toward the edge in the lower part of frame. SHOT 2, 3 to 5.5 seconds: Gimbal follow from behind as they walk hand in hand toward the guardrail, breath visible in the cold air. SHOT 3, 5.5 to 8 seconds: Slow crane up from behind them as they stop at the rail, the canyon unfolding below and sunlight sweeping across the walls. SHOT 4, 8 to 10 seconds: Close two-shot, faces lit warm. Leah turns to Ahmed and speaks; he laughs quietly and pulls her closer. DIALOGUE AND LIP-SYNC: Only Leah speaks, in Shot 4. Her mouth is clearly visible and lip-sync is exact. She says, soft and amazed: "Okay. The alarm was worth it." American accent, relaxed pace. SOUND: Vast, quiet open-air ambience: a light cold wind, gravel crunching under boots in Shot 2, a raven calling in the distance, the soft rustle of their jackets. A gentle acoustic guitar melody rises slowly under Shots 1 to 3 and fades under Leah's line. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound lands on the frame of the action that causes it, room acoustics match each space, and only the speaker's lips move when a line is spoken. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The canyon geology and the river stay consistent in position and scale between shots. - Their faces and jackets remain consistent from the drone shot to the close-up. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second widescreen food documentary moment at a busy taco truck on a Los Angeles street at night. The cook assembles al pastor tacos while talking to a customer. Four connected shots with consistent people, truck, and food, packed with sizzling detail and street energy. THE SUBJECT: Rosa, a Mexican American woman around 45, with dark hair in a low bun under a black hairnet, a black apron over a maroon T-shirt, and a small gold cross necklace. She works with a long knife and metal tongs. A customer, Tyler, a white American man around 25 in a gray hoodie, appears in Shot 4 at the window. Wardrobe stays identical across shots. SETTING: A white taco truck parked on a busy Los Angeles street at night, a hand-painted menu board beside the window, a vertical spit (trompo) of marinated pork topped with a pineapple, a steel flat-top with warming tortillas, bowls of chopped onion, cilantro, lime wedges, and red salsa. A lit sign above the window reads TACOS ROSA in red capitals. Cars pass behind with headlight streaks. LIGHTING: Bright warm fluorescent light inside the truck spilling out of the service window, mixed with the orange glow of streetlights and red taillight streaks behind. Steam and smoke catch the light. CAMERA: Shot 1 is a macro close-up on the trompo. Shot 2 is a top-down shot of the flat-top. Shot 3 is a slow push-in on the plate. Shot 4 is a medium shot from the street through the window. FRAMING (LANDSCAPE 16:9): Compose for a widescreen display. Use the width deliberately: place the subject on a rule-of-thirds line with the direction of movement or gaze leading into open space. Keep horizons level unless a shot calls for a tilt. Leave a comfortable margin on all sides so nothing important touches the frame edge. Do not letterbox or add black bars. SHOTS (total 10 seconds): SHOT 1, 0 to 2.5 seconds: Macro close-up as Rosa's knife shaves thin, caramelized slices of red al pastor pork off the spinning trompo; juices drip and a slice of pineapple is flicked off the top. SHOT 2, 2.5 to 5 seconds: Top-down shot of the flat-top as she lays out three small corn tortillas side by side, flips them with her fingers, then piles on the pork with tongs. SHOT 3, 5 to 7.5 seconds: Slow push-in on a paper boat of three tacos as she tops them with onion, cilantro, pineapple, and a drizzle of red salsa, then drops in a lime wedge. SHOT 4, 7.5 to 10 seconds: Medium shot from the street. Rosa leans out of the window under the TACOS ROSA sign and hands the tacos to Tyler, smiling as she speaks. DIALOGUE AND LIP-SYNC: Rosa is the only speaker, in Shot 4, mouth clearly visible and lip-sync exact. Warm and teasing, with a light Mexican American accent: "Three al pastor. You're gonna come back tomorrow." SOUND: Sizzling meat and dripping juices, the scrape of the knife, tortillas slapping on the flat-top, a crinkle of the paper boat, the truck's generator hum, passing cars, faint cumbia music playing from a radio inside the truck. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound lands on the frame of the action that causes it, room acoustics match each space, and only the speaker's lips move when a line is spoken. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - The sign reads TACOS ROSA exactly. - Food looks real with authentic char and texture; tortillas bend naturally. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

Prompt

A 10 second widescreen action sports sequence at a concrete skate park at golden hour. A skater lands a trick line and celebrates with a friend. Four connected shots designed to show fast motion and fine detail staying crisp, with consistent people and gear. THE SUBJECT: Nia, a Black American woman around 22, with shoulder-length box braids tied back, a loose white T-shirt, baggy faded blue jeans, a black beanie, and worn black and white skate shoes. Her skateboard has a natural maple deck with a bright yellow griptape strip and red wheels. Her friend Cole, a Korean American man around 23, in a green flannel shirt, filming on his phone. Wardrobe and board stay identical across shots. SETTING: An outdoor concrete skate park in Venice Beach, California: smooth gray bowls, a long metal handrail on a set of stairs, graffiti murals on a low wall, palm trees and the ocean in the background. A few other skaters sitting on the edge of the bowl. LIGHTING: Golden hour sun low over the ocean behind the park, warm backlight and long shadows on the concrete, sun flare through the palms. CAMERA: Shot 1 is a low follow shot at board height. Shot 2 is a fisheye close tracking shot. Shot 3 is a wide shot at real speed. Shot 4 is a handheld medium two-shot. FRAMING (LANDSCAPE 16:9): Compose for a widescreen display. Use the width deliberately: place the subject on a rule-of-thirds line with the direction of movement or gaze leading into open space. Keep horizons level unless a shot calls for a tilt. Leave a comfortable margin on all sides so nothing important touches the frame edge. Do not letterbox or add black bars. SHOTS (total 10 seconds): SHOT 1, 0 to 2.5 seconds: Low follow shot at board height as Nia pushes twice and picks up speed across the flat concrete toward the stairs, wheels humming. SHOT 2, 2.5 to 5 seconds: Fisheye tracking shot beside her as she ollies onto the handrail and grinds down it, sparks of sunlight glinting off the metal. SHOT 3, 5 to 7 seconds: Wide shot at real speed as she pops off the end of the rail and lands clean, rolling away while the skaters on the bowl edge cheer. SHOT 4, 7 to 10 seconds: Handheld two-shot as she rolls up to Cole, stops by stepping on the tail, and grins. He shows her the phone screen and speaks, and they bump fists. DIALOGUE AND LIP-SYNC: Only Cole speaks, in Shot 4, mouth clearly visible and lip-sync exact. Excited and loud: "First try! I got the whole thing!" American accent. SOUND: Board wheels rolling on smooth concrete, two push steps, the sharp pop of the ollie, the metallic grind along the rail, a solid landing clack, whoops and board taps from the watching skaters, distant waves and seagulls. IMAGE QUALITY AND TEXTURE: Treat this as footage captured on a real cinema or high-end phone camera, not a render. Keep true-to-life color with natural skin tones across every ethnicity, believable fabric texture and wrinkles, fine hair strands that move with the air, and small imperfections such as dust in light beams, fingerprints on glass where appropriate, or scuffs on worn surfaces. Exposure is balanced: highlights roll off smoothly and shadows keep detail. Motion blur matches a 180 degree shutter at the frame rate. Depth of field behaves like a real lens, with focus pulls that are smooth and purposeful. Backgrounds stay stable and do not shimmer, melt, or change layout between shots. AUDIO TO PICTURE SYNC: Every sound lands on the frame of the action that causes it, room acoustics match each space, and only the speaker's lips move when a line is spoken. RULES: - Photoreal, natural motion at real-world speed. No slow motion unless a shot asks for it, no morphing, no warping between frames. - Hands are anatomically correct with five fingers each, natural knuckles and nails, and they grip objects believably without clipping through them. - Faces, hair, wardrobe, props, and set dressing stay identical across every shot. The same person must look like the same person in every cut. - No on-screen captions, subtitles, lower thirds, logos, or watermarks. The only visible text is the text named in this prompt, spelled exactly as written. - Cuts happen only at the timecodes listed. Inside a shot the camera move is continuous. - Lighting direction and color temperature stay consistent between shots in the same location. - Board physics are real: feet stay planted during the grind and the board flips and lands believably. - The skateboard deck, griptape, and red wheels are identical in every shot. - Nia keeps her beanie on throughout; her braids move with the wind and the grind and never change length or color. - Audio is one continuous mix across the cuts. No background music unless the SOUND section asks for it. Spoken lines are clear, natural in pace, and never overlap each other.

100M+VIDEOS CREATED
14M+USERS WORLDWIDE
80+LANGUAGES SUPPORTED

Why creators choose LTX-2.5 Fast

Built for speed

LTX-2.5 Fast is the speed tier of Lightricks' LTX-2.5 family. Lightricks reports generation faster than real time on its own hardware, which makes it a good fit for drafts and high-volume social content.

Clips up to 20 seconds

The Fast tier reaches 20 seconds per generation, twice the 10 second limit of LTX-2.5 Pro. On Fliki you can choose any length from 6 to 20 seconds.

Native multi-shot

One generation can contain several connected shots. Lightricks says character, scene, lighting, visual style, and voice stay consistent across the cuts.

Synchronized audio

Audio and video are generated together, so ambience, effects, and speech follow the action on screen. Describe the sound in your prompt and get a finished clip.

First and last frame control

Start from an uploaded image, or set both the first and last frame and let the model fill the motion between them. Handy for product reveals and planned transitions.

Sharper detail in busy scenes

Lightricks' Diffusion Fidelity Rendering puts more compute into complex moments, and a diffusion video decoder keeps faces, fast motion, and on-screen text clearer.

720p and 1080p output

Render at 720p on the Standard plan for fast drafts, or 1080p on Premium for final delivery. Output runs at 25 fps.

Open-weights foundation

LTX-2.5 is released with open weights, so the model behind your Fliki clips is the same one the research and ComfyUI community builds on.

How it works

How to generate a video with LTX-2.5 Fast

Getting a production-ready video out of LTX-2.5 Fast takes under a minute inside Fliki. Follow these six steps.

Fliki prompt input with a multi-shot text-to-video description for LTX-2.5 Fast
Step 1

Write your prompt

Open Fliki and describe the scene in plain language. Include the subject, setting, camera move, lighting, and sound. For multi-shot clips, list each shot in order so the model knows where to cut.

Fliki model selector with LTX-2.5 Fast chosen for AI video generation
Step 2

Select LTX-2.5 Fast as your model

Open the model selector and choose LTX-2.5 Fast. Fliki routes your prompt to Lightricks' model with no GPUs, weights, or setup on your side.

Choose 16:9 or 9:16 aspect ratio for LTX-2.5 Fast video generation on Fliki
Step 3

Pick your aspect ratio

Choose 16:9 for YouTube and landscape web, or 9:16 for TikTok, Reels, and Shorts. LTX-2.5 Fast composes both natively.

Set a 6 to 20 second duration for LTX-2.5 Fast on Fliki
Step 4

Set the duration

Pick a length from 6 to 20 seconds in 2 second steps. Longer clips give multi-shot sequences room to breathe; shorter clips are quicker to iterate on.

Upload first and last frame images for LTX-2.5 Fast image-to-video on Fliki
Step 5

Add a first or last frame (optional)

Upload a still to start the clip from, or a first and last frame to control where the shot begins and ends. Useful for product shots and matching a brand look.

Pick 720p or 1080p and generate a video with LTX-2.5 Fast on Fliki
Step 6

Select resolution and generate

Pick 720p for quick turnaround or 1080p for sharper detail, then hit Generate. The clip arrives with audio, ready to preview, download, or add to a Fliki project.

AI MODEL GALLERY

Built on the best AI models - ready inside Fliki

Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.

LTX-2.5 Fast FAQ

Frequently asked questions

Everything you need to know about generating with LTX-2.5 Fast inside Fliki.

Still curious?

Try Fliki free in your browser, no credit card required.

Start free
LTX-2.5 Fast · Free forever plan

Generate your next video with LTX-2.5 Fast.

Fast clips up to 20 seconds with synchronized audio and multi-shot scenes from Lightricks' open-weights model. Free to start, no credit card required.

Generate your first video free

Free forever plan · No credit card required · Cancel anytime